You write the code.
You write the tests.
Tests validate your work.
bottleneck: developer writing speed
AI writes the code — fast, confident,
syntactically correct, sometimes subtly wrong.
bottleneck: verification infrastructure
The developer's role shifts — from writing every line to defining intent and verifying correctness. The agent implements; you set the bar it must clear.
Developer drives all three steps.
TDD is a development loop.
TDD shifts from a development discipline to a
design communication tool — authoring intent for both humans and machines.
@Nested
@DisplayName("API retry policy")
class ApiRetryPolicyTest {
@Test @DisplayName("retries transient errors up to 3 times")
void retriesTransientErrors() { ...}
@Test @DisplayName("does not retry on authentication failures")
void noRetryOnAuthFailure() { ...}
}
@Mock KubernetesClient client;
@Test void deletesPod() {
service.delete("foo");
// agent renames pods()→resources()
// agent rewrites this mock in lockstep
// nothing pins the actual contract
verify(client.pods()
.withName("foo")).delete();
}Test and impl travel together. No guardrail survives the refactor.
@QuarkusTest @WithKubernetesTestServer
class PodBehaviorIT {
@Test
@DisplayName("DELETE removes the pod")
void removesPodFromCluster() {
k8s.pods().resource(pod).create();
given().delete("/api/v1/pods/foo")
.then().statusCode(204);
assertThat(k8s.pods()
.withName("foo").get()).isNull();
}
}Zero test edits after agent refactor. The outcome is the contract.
A 35-minute CI pipeline turns an AI agent into an expensive queue.
The agent spawns dozens of PRs per hour — they all wait in the same line. Your CI is the rate-limiter on your entire agentic workflow.
fabric8io/kubernetes-client
★ 3.7k · >1M downloads/month · 11 years
75+ Maven modules · multiple HTTP backends
(OkHttp · Vert.x · JDK HttpClient · Jetty)
Powers: Quarkus · Spring Cloud K8s · Apache Flink
30–40 minutes per PR.
Sequential jobs. Flaky tests retried 2–5× per run.
Every contributor.
Every PR.
Every day.
The CI became a throughput cap on all velocity — including AI-assisted work.
| module / job | before | after | cut |
|---|---|---|---|
| kubernetes-tests (surefire) | 266s | 109s | −59% |
| kubernetes-client-api (surefire) | 84s | 31s | −62% |
| kubernetes-client (surefire) | 153s | 43s | −72% |
| Windows build (wall-clock) | ~40 min (flaky, retried 2–5×) | ~22 min (stable) | −45% |
ExecWebSocketListener.onError dropped the process exit code on peer-close race. cleanUpOnce() shut down SerialExecutor, discarding the queued handleExitStatus. Callers saw silent failures.
ExecListener.onClose/onFailure lost on close-handshake race. Callers blocked on a latch indefinitely on transport drop. No timeout, no error signal.
SharedProcessor.distribute() threw NPE when called after informer.stop() — SerialExecutor.execute() invoked post-shutdown.
@EnableKubernetesMockClient(https=false) · Fabric8 KubernetesMockServer@EnableKubeAPIServer · Fabric8 kube-api-test (envtest)@QuarkusTest @WithKubernetesTestServerTestcontainers + KinD · real container runtime, real kubeletKubernetesClientBuilder smoke tests · pre-release only@Nested
@DisplayName("PodService")
class PodServiceTest {
@Nested @DisplayName (“when deleting a pod”)
class WhenDeletingAPod {
@Test
@DisplayName("removes it from the cluster")
void removesPodFromCluster() { ... }
@Test
@DisplayName("returns 404 for unknown pod")
void returnsNotFoundForUnknownPod() { ... }
} }
An AI agent given this test hierarchy understands exactly what to implement — not from a prose comment, but from code that runs on every commit.
Text docs are advisory.
Tests push back.
// a failing @DisplayName reads like a sentence top to bottom
Fix one test this week — rename it to describe what it guards, not how it works.
For Kubernetes-facing code, KubernetesMockServer gives you ~1s/class — no cluster, no Docker, no waiting.
An agent running a 1-second test loop can make 36 attempts in the time it waits for one CI run.
E2E / KinD on PR label or main branch only.
Integration tests in a parallel matrix job.
Smoke tests pre-release, not every commit.
Goal: sub-10-min critical path on every PR.
Each flaky test is a retry cost multiplied across every contributor's every PR.
Under AI-driven velocity, flake debt compounds fast.
Turn one flaky test into a real, deterministic failing test this week.
Tests define intent for both humans and AI agents. The spec is always current, always enforced.
Fast feedback loop = more agent cycles per hour. Mistakes caught in seconds, not 35 minutes later.
Review scope, architecture, correctness — not syntax. You keep the judgment. AI takes the grind.
More iterations → more edge cases caught → higher confidence in AI output → agents trusted with more.
When AI agents write your code, your tests define what should be built. TDD shifts from a development discipline to a design communication tool — you're authoring intent for both humans and machines.
AI agents iterate at machine speed, but only if your feedback loop keeps up. A 35-minute pipeline turns an agent into an expensive queue. Sub-10-minute pipelines are what unlock the real gains — as we proved on fabric8.
AI-generated code is confident, syntactically correct, and sometimes subtly wrong. A comprehensive, fast test suite is the only scalable way to maintain confidence in code you didn't write line-by-line.