←
~/dev/techgenie-2026 · talk.md 01

Your Tests Are Your Spec

// Why TDD and Fast CI/CD Are the Real Enablers of Agentic Software Development
Ashish Thakur
Senior Software Engineer · Red Hat
TechGENIE 2026 · 09 September 2026
~/dev/techgenie-2026 · whoami.md 02

About Me

Ashish Thakur
Ashish Thakur
Senior Software Engineer at Red Hat. Open source maintainer & conference speaker.

Pinned
Java client for Kubernetes & OpenShift — 75+ Maven modules, multiple HTTP backends (OkHttp · Vert.x · JDK · Jetty)
Build & deploy Java apps to Kubernetes/OpenShift — Maven & Gradle plugins, zero-config Docker image build
Highlights
★ 3.7kstars · fabric8 kubernetes client
>1M / moMaven downloads across all modules
11 yrsopen source maintenance & contribution
31flaky tests fixed → 5 real production bugs found
powersQuarkus · Spring Cloud Kubernetes · Apache Flink · JOSDK
~/dev/techgenie-2026 · Act01.java 03
[INFO] Building io.techgenie:talk 1.0.0-SNAPSHOT
▶act-01 · the-shift ................................BUILDING ...
·act-02 · tests-are-your-spec ......................PENDING
·act-03 · fast-feedback-loops ......................PENDING
·act-04 · patterns .....................................PENDING
// act 01

The *Shift*

~/dev/techgenie-2026 · TheShift.java act 01 · 04

AI writes code. You verify it.

✗  Old World

You write the code.
You write the tests.
Tests validate your work.

bottleneck: developer writing speed

▶  Agentic World

AI writes the code — fast, confident,
syntactically correct, sometimes subtly wrong.

bottleneck: verification infrastructure

The developer's role shifts — from writing every line to defining intent and verifying correctness. The agent implements; you set the bar it must clear.

without verification infrastructure, AI-assisted development is AI-assisted gambling
~/dev/techgenie-2026 · TheAmplifier.java act 01 · 05

AI amplifies what’s already there

✗  Poor hygiene → slop, faster
  • Huge PRs with no baby steps
  • No behavioral tests or spec docs
  • Vague issues, no acceptance criteria
  • Slow, flaky CI pipeline
  • Problems hidden for months → exposed in hours
✓  Good hygiene → acceleration
  • Small, atomic, reviewable PRs
  • Behavioral tests that enforce intent
  • Well-defined issues with acceptance criteria
  • Fast, stable CI with real feedback
  • Guardrails catch AI mistakes before prod
"Legacy is not age. Legacy is unsafe change." — an AI-assisted codebase without tests becomes legacy fast
~/dev/techgenie-2026 · Act02.java 06
[INFO] Building io.techgenie:talk 1.0.0-SNAPSHOT
✓act-01 · the-shift ................................SUCCESS [03:30]
▶act-02 · tests-are-your-spec ......................BUILDING ...
·act-03 · fast-feedback-loops ......................PENDING
·act-04 · patterns .....................................PENDING
// act 02

Tests Are *Your Spec*

~/dev/techgenie-2026 · TDDThenVsNow.java act 02 · 07

TDD: the practice that changed its meaning

Traditional TDD
①Write a failing test
②Write code to make it pass
③Refactor

Developer drives all three steps.

TDD is a development loop.

Agentic TDD
①Write a failing test(you define intent)
②AI writes code to pass it
③CI validates it

TDD shifts from a development discipline to a

design communication tool — authoring intent for both humans and machines.

~/dev/techgenie-2026 · Spec.java act 02 · 08
// specs
explain intent
// tests
enforce it
A prose spec is advisory.  A test is enforced on every commit.
When the spec drifts, no one notices.  When the test fails, the build fails.
~/dev/techgenie-2026 · TestAsSpec.java act 02 · 09

A failing test is a machine-readable spec

@Nested
@DisplayName("API retry policy")
class ApiRetryPolicyTest {

  @Test @DisplayName("retries transient errors up to 3 times")
  void retriesTransientErrors() { ...}

  @Test @DisplayName("does not retry on authentication failures")
  void noRetryOnAuthFailure() { ...}
}
✗ BUILD FAILED — ApiRetryPolicyTest
  API retry policy
    ✗ does not retry on authentication failures
      expected: 1 attempt · actual: 3 attempts
the failing test reads like a sentence. That sentence is your spec. An AI knows exactly what to implement.
~/dev/techgenie-2026 · BlackBoxVsWhiteBox.java act 02 · 10

Write tests that survive AI refactors

✗  White-box — mocks internals
@Mock KubernetesClient client;
@Test void deletesPod() {
    service.delete("foo");
    // agent renames pods()→resources()
    // agent rewrites this mock in lockstep
    // nothing pins the actual contract
    verify(client.pods()
           .withName("foo")).delete();
}

Test and impl travel together. No guardrail survives the refactor.

✓  Black-box — asserts outcomes
@QuarkusTest @WithKubernetesTestServer
class PodBehaviorIT {
  @Test
  @DisplayName("DELETE removes the pod")
  void removesPodFromCluster() {
    k8s.pods().resource(pod).create();
    given().delete("/api/v1/pods/foo")
           .then().statusCode(204);
    assertThat(k8s.pods()
        .withName("foo").get()).isNull();
  }
}

Zero test edits after agent refactor. The outcome is the contract.

~/dev/techgenie-2026 · Act03.java 10
[INFO] Building io.techgenie:talk 1.0.0-SNAPSHOT
✓act-01 · the-shift ................................SUCCESS [03:30]
✓act-02 · tests-are-your-spec ......................SUCCESS [04:20]
▶act-03 · fast-feedback-loops ......................BUILDING ...
·act-04 · patterns .....................................PENDING
// act 03

Fast *Feedback Loops*

~/dev/techgenie-2026 · SpeedMismatch.java act 03 · 11

AI moves at machine speed. CI doesn’t.

fabric8io/kubernetes-client/Pull requests ·Open (4) — AI-generated in the last 10 minutes
Open #842 · feat: add retry logic for transient API failures ⏳ CI queued · 34m ago
Open #843 · fix: handle null namespace in informer shutdown path ● CI running · 28m elapsed
Open #844 · test: add missing coverage for SSL body-stream edge case ● CI running · 12m elapsed
Open #845 · refactor: extract shared serial executor helper ⏳ CI queued · just now

A 35-minute CI pipeline turns an AI agent into an expensive queue.
The agent spawns dozens of PRs per hour — they all wait in the same line. Your CI is the rate-limiter on your entire agentic workflow.

fast CI/CD is a multiplier, not a luxury — sub-10-minute pipelines unlock the real productivity gains of agentic development
~/dev/techgenie-2026 · Fabric8Project.java act 03 · 12

Case study: fabric8 Kubernetes Client

Project

fabric8io/kubernetes-client

★ 3.7k  ·  >1M downloads/month  ·  11 years

75+ Maven modules · multiple HTTP backends
(OkHttp · Vert.x · JDK HttpClient · Jetty)

Powers: Quarkus · Spring Cloud K8s · Apache Flink

The CI problem

30–40 minutes per PR.

Sequential jobs. Flaky tests retried 2–5× per run.

Every contributor.
Every PR.
Every day.

The CI became a throughput cap on all velocity — including AI-assisted work.

~/dev/techgenie-2026 · CIPipeline.java act 03 · 13

Before and after: the pipeline transformation

✗  Before — sequential, slow, flaky
fabric8io/kubernetes-client/Actions·Run #1823 · PR #7800
✗ CI — 38m 42s · 2 failed jobs sequential · no cache
✗unit-tests18m 23s× 2 retries
○unit-tests (retry 1)18m 41sstill flaky
✓build8m 11s
✓style-check5m 04s
✓sonar7m 04s
✓  What we changed
  • KubernetesMockServer — no real cluster in unit tests
    ~50s/class → ~1s/class
  • Parallel matrix — Java versions & HTTP backends run concurrently
  • Maven dep caching — eliminated cold-start download cost
  • Selective E2E — KinD cluster tests on PR label or main only
~/dev/techgenie-2026 · Results.java act 03 · 14

Real numbers — ~40 PRs across 4 epics

module / jobbeforeaftercut
kubernetes-tests (surefire)266s109s−59%
kubernetes-client-api (surefire)84s31s−62%
kubernetes-client (surefire)153s43s−72%
Windows build (wall-clock)~40 min (flaky, retried 2–5×)~22 min (stable)−45%
−45% billable CI minutes
per PR
~10 min critical-path CI
(happy path)
$0.005/min runner cost after
macOS → ARM64
fabric8io/kubernetes-client/Pull requests
Merged #7812 · test: use ephemeral ports in httpclient-vertx SSL/body-stream tests
✓  All checks have passed
✓CI / unit-tests (Java 21 · OkHttp)1m 49s
✓CI / unit-tests (Java 17 · Vert.x)1m 31s
✓CI / unit-tests (Java 21 · JDK HttpClient)1m 22s
✓CI / build · sonar48s
~/dev/techgenie-2026 · HiddenBugs.java act 03 · 15

Fixing flaky tests surfaced real production bugs

fabric8io/kubernetes-client/Actions·31 flaky issues closed in ~8 weeks · 5+ were real production bugs
✗ CI — tests started failing deterministically after flake fix not flaky · not noise · real bugs
✗ExecWebSocketListenerTest · testHandleExitStatusexit code dropped on peer-close race → #7700
✗ExecListenerTest · testOnCloseRacecallers block forever on transport drop → #7779
✗SharedInformerTest · testStopRaceNPE in SerialExecutor post-stop → #7716

#7700 — Exit code silently dropped

ExecWebSocketListener.onError dropped the process exit code on peer-close race. cleanUpOnce() shut down SerialExecutor, discarding the queued handleExitStatus. Callers saw silent failures.

#7779 — Callers blocked forever

ExecListener.onClose/onFailure lost on close-handshake race. Callers blocked on a latch indefinitely on transport drop. No timeout, no error signal.

#7716 — NPE after stop()

SharedProcessor.distribute() threw NPE when called after informer.stop() — SerialExecutor.execute() invoked post-shutdown.

when AI makes everything move fast, tech debt shows up as a throughput cap — and flaky tests are the canary
~/dev/techgenie-2026 · TheLegacyLoop.java act 03 · 16

The loop you remember from 2010 is the loop your agent is trapped in

☠ legacy WebLogic 10.3 · app.ear · 2010 ⏱ ~30 minutes
[14:03:11] <Notice> <WebLogicServer> Starting WebLogic Server…
[14:05:42] <JDBC> pool 'AppPool' created, awaiting datasource…
[14:08:19] <Deployer> Deploying application 'app.ear'…
[14:14:07] <EJB> Initializing 312 session beans…
[14:21:55] <JNDI> Re-binding 187 names into context…
[14:28:34] <Warning> Slow startup. Re-deploy after each change.
[14:31:02] Application available at /app
an agent is only as fast as the feedback loop — the loop is the bottleneck
~/dev/techgenie-2026 · FeedbackPyramid.java act 03 · 17

The feedback pyramid: many fast, few slow

~1s / class
In-process mock · @EnableKubernetesMockClient(https=false) · Fabric8 KubernetesMockServer
every save
~3s / class
Real apiserver · @EnableKubeAPIServer · Fabric8 kube-api-test (envtest)
every save
~5s / class
App booted in-process · @QuarkusTest @WithKubernetesTestServer
every commit
~50s / class
KinD cluster · Testcontainers + KinD · real container runtime, real kubelet
every PR
~10 min / suite
Real dev cluster · KubernetesClientBuilder smoke tests · pre-release only
pre-release
many tests at the top, few at the bottom — agent iterates where it's cheap
~/dev/techgenie-2026 · Act04.java 18
[INFO] Building io.techgenie:talk 1.0.0-SNAPSHOT
✓act-01 · the-shift ................................SUCCESS [03:30]
✓act-02 · tests-are-your-spec ......................SUCCESS [04:20]
✓act-03 · fast-feedback-loops ......................SUCCESS [06:10]
▶act-04 · patterns .....................................BUILDING ...
// act 04

*Patterns* to Start Today

~/dev/techgenie-2026 · Pattern1.java act 04 · 19

Pattern 1: write tests that communicate intent

Use @DisplayName as your spec language
@Nested
@DisplayName("PodService")
class PodServiceTest {

@Nested @DisplayName (“when deleting a pod”) class WhenDeletingAPod {

@Test
@DisplayName("removes it from the cluster")
void removesPodFromCluster() { ... }

@Test
@DisplayName("returns 404 for unknown pod")
void returnsNotFoundForUnknownPod() { ... }

} }

The AI reads this as a contract

An AI agent given this test hierarchy understands exactly what to implement — not from a prose comment, but from code that runs on every commit.

Text docs are advisory.
Tests push back.

// a failing @DisplayName reads like a sentence top to bottom

Fix one test this week — rename it to describe what it guards, not how it works.

~/dev/techgenie-2026 · Pattern2.java act 04 · 20

Pattern 2: build your fast feedback loop first

Invest in in-process mocks

For Kubernetes-facing code, KubernetesMockServer gives you ~1s/class — no cluster, no Docker, no waiting.

An agent running a 1-second test loop can make 36 attempts in the time it waits for one CI run.

Gate slow tests selectively

E2E / KinD on PR label or main branch only.

Integration tests in a parallel matrix job.

Smoke tests pre-release, not every commit.

Goal: sub-10-min critical path on every PR.

Treat flaky tests as bugs

Each flaky test is a retry cost multiplied across every contributor's every PR.

Under AI-driven velocity, flake debt compounds fast.

Turn one flaky test into a real, deterministic failing test this week.

~/dev/techgenie-2026 · CompoundingAdvantage.java act 04 · 21

The compounding advantage

1

Write behavioral tests

Tests define intent for both humans and AI agents. The spec is always current, always enforced.

2

AI iterates faster

Fast feedback loop = more agent cycles per hour. Mistakes caught in seconds, not 35 minutes later.

4

Developer gains capacity

Review scope, architecture, correctness — not syntax. You keep the judgment. AI takes the grind.

3

Coverage compounds

More iterations → more edge cases caught → higher confidence in AI output → agents trusted with more.

every improvement lowers the cost of the next change. Teams that build this infrastructure first compound their advantage.
~/dev/techgenie-2026 · Takeaways.java act 04 · 22

Key takeaways

01
Tests are the new specification language.

When AI agents write your code, your tests define what should be built. TDD shifts from a development discipline to a design communication tool — you're authoring intent for both humans and machines.

02
Fast CI/CD is a multiplier, not a luxury.

AI agents iterate at machine speed, but only if your feedback loop keeps up. A 35-minute pipeline turns an agent into an expensive queue. Sub-10-minute pipelines are what unlock the real gains — as we proved on fabric8.

03
Trust requires verification infrastructure.

AI-generated code is confident, syntactically correct, and sometimes subtly wrong. A comprehensive, fast test suite is the only scalable way to maintain confidence in code you didn't write line-by-line.

~/dev/techgenie-2026 · Resources.md 23

Resources

AI takes the repetitive grind. You keep the judgment.
~/dev/techgenie-2026 · Thanks.java 24

Thanks!

Ashish Thakur
Senior Software Engineer · Red Hat
[INFO] ─────────────────────────────────────────────
[INFO] BUILD SUCCESS
[INFO] Total time: 15:00 min
[INFO] ─────────────────────────────────────────────