DevelopmentAI Engineering

Specs, Traffic, or Agents? Three Ways to Test APIs With AI (2026)

Three ways AI generates API tests now: from your OpenAPI spec (Schemathesis), from recorded traffic (Keploy), or via workspace agents (Postman Agent Mode, Kiro). A worked endpoint with three planted bugs shows what each approach catches and what each one structurally misses, plus how to layer all three in CI.

Toolbit AI - Team
14 min read
Specs, Traffic, or Agents? Three Ways to Test APIs With AI (2026)

Your API has a spec. It also has behavior the spec doesn't know about.

Every API team eventually discovers this the hard way. The OpenAPI document says a field is an integer with a minimum of 1, so you test 1, 2, and 50. Then a customer submits an order for nine hundred million units, the inventory service overflows a signed int somewhere down the call chain, and your API returns a 500 with a stack trace leaking into the response body. The spec was accurate. The spec was also useless in that moment, because no generated test probed what the spec allowed, only what a human thought to write down.

That gap between documented behavior and actual behavior is exactly where the current wave of AI API testing tools aims. But they don't all aim at the same target. In September 2026, the tools worth knowing about cluster into three genuinely different approaches: generating tests from your specification, deriving tests from recorded production traffic, and pointing an AI agent at your workspace to explore what it observes. Each one catches a different class of bug, and each one is blind to a different class of bug. Picking wrong isn't expensive in dollars, it's expensive in the false confidence of a green test suite that never touched the input that mattered.

Here's what each approach actually finds, worked through a single concrete endpoint, with an honest account of what each one misses.


The three ways tests get generated now

Quick orientation before the details:

  • Spec-first generation. The tool reads your OpenAPI or GraphQL schema and mechanically derives test cases from its constraints: types, ranges, required fields, enums. Schemathesis is the strongest open source example, using property-based testing to probe boundary values and constraint violations automatically.
  • Traffic-derived tests. The tool records real API calls as they happen, then converts those recordings into replayable tests with assertions and mocks for downstream dependencies. Keploy is the established open source player here, accepting OpenAPI specs, Postman collections, or raw recordings.
  • Agent-assisted exploratory. An AI agent sits inside your testing workspace, watches real requests and responses, and writes tests from what it sees, plus what it can infer. Postman's Agent Mode does this inside Postman. AWS's Kiro does something adjacent from the other direction: it writes requirements as formal specs first, then checks the implementation against them with property-based tests.

The interesting question isn't which is best. It's what each one catches, and what each one structurally cannot.


A worked example: three planted bugs in one endpoint

To make this concrete, take a single endpoint. A simplified POST /orders with a request body schema like this:

POST /orders
requestBody:
  content:
    application/json:
      schema:
        type: object
        required: [items]
        properties:
          items:
            type: array
            minItems: 1
            items:
              type: object
              required: [sku, quantity]
              properties:
                sku: { type: string }
                quantity: { type: integer, minimum: 1 }
responses:
  "201": { description: Order created }
  "400": { description: Validation error }
  "401": { description: Unauthorized }

Now plant three bugs that the spec cannot describe, because all three live in undocumented behavior:

  1. The overflow. Any integer greater than or equal to 1 is spec-valid, but a quantity above 900 million overflows the inventory service and returns a 500.
  2. The undocumented conflict. A duplicate Idempotency-Key header returns a 409. The spec documents 201, 400, and 401 for this endpoint. It says nothing about 409.
  3. The silent coercion. A quantity of 2.5 (a float, spec-invalid) is silently truncated to 2 and accepted with a 201, instead of being rejected with a 400.
Matrix of three planted bugs against spec-first, traffic-derived and agent-assisted testing approaches

Watch what each approach does with this endpoint.


Spec-first: exhaustive on the contract, blind beyond it

Point Schemathesis at this spec and run it. Property-based testing means it doesn't need you to enumerate cases: it reads minimum: 1 and generates 1, 2, values at the boundary, values deep in the range, and deliberately invalid inputs: zero, negative numbers, floats, strings, nulls. Then it checks universal properties against every response, things like "the server must never return a 5xx for a spec-valid request" and "every response must conform to the schema declared for that status code."

Against our three planted bugs:

  • The overflow gets caught immediately. A large spec-valid integer triggering a 500 is a direct violation of the core not_a_server_error check. This is the bug class Schemathesis exists for. Better, because it's built on the Hypothesis library, when it finds the failure it shrinks it: you don't get "some huge number broke it," you get the minimal failing value, plus a ready-to-paste curl command for the bug report.
  • The 409 gets caught as a conformance failure. The status_code_conformance check flags it: the server returned a status code the schema never documented. Schemathesis won't know it's about idempotency, but it will tell you exactly which request produced an undocumented response, which is the lead you need.
  • The silent coercion gets caught, conditionally. Its negative-data checks flag that a spec-invalid float was accepted. Whether that registers as a bug or noise depends on how you configure strictness, but the signal is there in the report.

What it misses is everything the schema doesn't say. Schemathesis knows nothing about authorization rules beyond what's in securitySchemes, nothing about side effects, nothing about whether that 409 represents correct business logic or a bug. Its ceiling is the quality of your spec, full stop. A spec that documents quantity as an unbounded integer but should have said maximum: 999 will produce tests that faithfully exercise a wrong contract. The vendor reports that production schemas typically surface five to fifteen issues on a first run, and the reason is almost always that the spec drifted from the implementation, not that the implementation drifted from the spec.

The maintenance story is the best of the three approaches: tests regenerate from the schema on every run, so a changed endpoint means changed tests with no per-endpoint editing. Schemathesis supports OpenAPI 2.0 through 3.2 plus GraphQL, runs from a schemathesis.toml config with no code required to start, and ships a GitHub Action and JUnit/Allure output for CI gating. If your spec is a live, honest document, this approach scales across your whole API surface for near-zero ongoing effort.


Traffic-derived: it tests reality, but only the reality that happened

Keploy starts from the opposite end. Instead of the contract, it records actual API traffic, either through its Chrome extension while you exercise the app, through a local agent for firewalled services, or from an OpenAPI spec or Postman collection you paste in. Each recorded interaction becomes a test case with assertions derived from the real response, plus mocks for every dependency the call touched, so the test can replay in an isolated CI sandbox without a database or third-party service being live.

Against our three bugs, with a recording window of normal usage:

  • The overflow is invisible. No real customer ordered 900 million units during recording. If it never happened, it never became a test. This is the structural blind spot of the approach: it covers observed behavior exhaustively and unobserved behavior not at all.
  • The 409 only appears if someone double-clicked. If some user retried a submit during the recording window, you get a captured 409 and a regression test for free, including the mock of whatever the idempotency store did. If nobody did, the behavior stays dark.
  • The coercion is the same story. If some client with a sloppy serializer ever sent 2.5, Keploy pinned it. Otherwise, nothing.

That sounds like a harsh verdict, so here's the counterweight: for regressions in behavior that does occur, traffic-derived tests are unmatched. A recorded checkout flow, replayed on every commit with mocked dependencies, catches integration breakage that no spec-based tool can, because the spec doesn't describe integrations at all. Keploy also generates full lifecycle flows (create, mutate, delete) rather than isolated requests, deduplicates similar cases, and detects flaky tests by re-running them. When the API changes shape, it self-heals minor differences rather than dumping a wall of red on your CI.

The maintenance burden sits in the middle of the three approaches. Recordings go stale the way fixtures do, but re-recording from fresh traffic replaces a whole batch at once, which is a different economics from editing test code by hand. The infrastructure cost is real though: you're running a recording layer, managing mocks, and keeping the sandbox honest.


Agent-assisted: it can guess the edge cases, but you have to check its homework

This is the newest and least mechanical of the three. Postman's Agent Mode, which replaced the older Postbot assistant in Postman's v12 platform, is an AI agent that lives inside your workspace with access to your collections, environment variables, response data, and specs. You send a request, get a response, and describe what you want tested in plain language. It writes standard Postman test scripts: plain JavaScript that runs in Collection Runner, the Postman CLI, and CI without any special runtime.

What makes it different from template-based generators is that it reasons over observed data. Show it a user response and it notices the usr_ prefix on IDs and writes a regex assertion for it. It sees email-shaped fields and validates the format. Point it at a 401 and it traces the failure back through your environment variables and pre-request scripts to tell you which token refresh flow is broken, which is diagnosis, not just test generation.

Against our three planted bugs, the honest answer is: it depends on what you ask. Postman's own practical guide is unusually candid about this. The agent focuses on what it can observe, so if you want tests for invalid inputs, missing fields, and unauthorized access, you have to ask for them explicitly. And it can generate confidently wrong assertions: their example is an assertion that a name field equals "Test Developer" (the literal test data) when you wanted a type check. Every generated test needs review before it lands in the collection. The agent does ask for approval before modifying collections, which keeps the blast radius contained.

What an agent does that the other two approaches genuinely cannot: it brings domain intuition. A well-prompted agent will think of the duplicate idempotency key, the float where an integer belongs, the absurdly large quantity, because those are patterns it has seen across thousands of APIs, not just yours. It can then send those requests and observe the undocumented 409 and the silent truncation directly. But it does this on request, not systematically, and it can invent an expectation that's wrong. You get exploratory reach in exchange for verification burden.

Kiro, AWS's spec-driven IDE, attacks the same territory from a different angle that's worth knowing about. Kiro's bet is that the fix isn't better test generation, it's better requirements: prompts become formal specs (requirements with EARS-notation acceptance criteria, design docs, sequenced tasks) before any code is written, and its Correctness feature checks requirements for contradictions and gaps using automated reasoning, then runs property-based tests against the implementation. Its hooks system adds maintenance automation: when you save a component or modify an API endpoint, event-driven agent hooks update the associated test or documentation automatically. It's an IDE-first workflow rather than a testing tool you point at an existing API, and it's priced accordingly: a perpetual free tier with 50 credits and access to Claude Sonnet 4.5 plus open weight models, then paid tiers from $20 per user per month (1,000 credits) up to $200 per month (10,000 credits) as of late 2026.

Agent-generated tests carry the heaviest maintenance burden of the three, because generated scripts are just code and rot like code. The mitigation is that the same agent class that wrote them can repair them: Postman's Agent Mode diagnoses failing scripts when asked, and Kiro hooks keep test files updated on save. But "the AI maintains the AI's tests" is a workflow you supervise, not one you forget about.


What each approach catches, side by side

Bug classSpec-first (Schemathesis)Traffic-derived (Keploy)Agent-assisted (Postman, Kiro)
Crash on spec-valid edge input (our overflow)Catches it, shrunk to minimal reproMisses it unless traffic contained itCatches it if prompted to probe ranges
Undocumented status codes (our 409)Flags as conformance failureCaptures it only if it occurredDiscovers it if it guesses the scenario
Invalid input accepted (our float coercion)Flags via negative-data checksOnly if a client did itCatches it if asked for negative tests
Business logic errors (wrong price, bad state transition)Cannot see itPins whatever actually happenedReasonable guesses, needs human judgment
Integration regressions (downstream breakage)Blind, spec doesn't describe depsBest coverage, replay with mocksPartial, depends on workspace context
Schema drift (response stops matching spec)Core strength, every response checkedDrift breaks recordings until re-recordCan spot it in diffs if asked
Maintenance costNear zero, regenerates from specMedium, re-record on API changeHighest per test, but agents can repair

The maintenance question, honestly

The reason teams abandon generated test suites isn't generation quality, it's that six months later the suite is red for reasons nobody has time to diagnose, so it gets skipped in CI, so it becomes decoration.

Spec-first generation largely escapes this trap because the schema is the single source of truth for both the API and its tests. Change the schema, and the next run tests the new contract automatically. Schemathesis also supports baseline files so CI fails only on new failures, which is the difference between a usable gate and a wall of pre-existing noise. The catch is that all this discipline collapses if the spec rots, and spec rot is the failure mode nobody budgets for.

Traffic-derived suites degrade differently: they don't produce false failures, they produce recordings of an API that no longer exists. Keploy's answer is self-healing and re-recording, which works well for minor changes and buys you a batch operation instead of line editing for major ones. The discipline you need is scheduled re-recording, the same way you'd refresh fixtures.

Agent-generated suites are ordinary code and need the ordinary care: review on merge, repair on breakage, deletion when the behavior they pin goes away. What changes is that the marginal cost of repair drops, because you can describe the fix rather than write it. What doesn't change is that someone still has to judge whether the fix was right. If you want to go deeper on giving agents the right context in the first place, our piece on the context engineering skill that's replacing prompt tweaking covers the discipline that decides whether your agent writes good tests or plausible-looking ones.


How to actually combine them

None of these approaches is a complete answer, which is why the mature setups in 2026 layer them:

  1. Run spec-first fuzzing in CI as the floor. One Schemathesis run against your live schema catches the 500s, undocumented status codes, and schema drift that hand-written tests structurally miss. Start with a baseline file so you're only failing on new issues, and tighten it as you fix. It's open source, config-driven, and costs nothing per endpoint.
  2. Record your critical flows once, replay forever. Use Keploy (or any record-replay tool) for the handful of multi-service flows where breakage is expensive: checkout, auth, provisioning. These tests earn their keep because they pin real integration behavior with real mocks, something no spec can express.
  3. Spend agent time where judgment is scarce. Don't ask an agent to generate a thousand assertions you'll never review. Ask it to explore: "what edge cases would break this endpoint," then verify the interesting guesses by hand or promote them into the spec as constraints. Review everything it writes; Postman's own documentation says so, and it's right.
Three-step CI layering flow: spec-first fuzzing, replay real flows, agents explore

If your team is already living in agentic tooling, the plumbing increasingly matters: Postman ships an MCP server for AI tool integration, and Kiro supports MCP natively, so your test tooling can be driven from the same agent stack you use for coding. Our developer-focused overview of the Model Context Protocol explains how that layer fits together, and if you're evaluating agentic IDEs more broadly, our comparison of the AI programming tools developers actually adopted puts Kiro's spec-driven approach in context.

One warning before you wire everything together: don't let an agent run exploratory tests against production. Property-based fuzzing and agent-probed edge cases belong in staging, with rate limits, against data you can poison. The first tool that discovers your overflow should not be the same request that also corrupted your inventory table.


Two questions people actually ask

Do AI-generated API tests replace hand-written ones? No, and the vendors mostly don't claim they do. Generation covers the mechanical mass: happy paths, documented error codes, boundary values, contract conformance. Hand-written tests remain the only place business logic gets asserted, because neither a schema nor a recording knows that a 409 should only happen for orders above a certain amount. The realistic win is that generated tests free your team's writing time for the tests only humans can design.

Is spec-first or traffic-derived better if I can only pick one? Spec-first, if you have a spec that's even 80 percent honest, because it probes inputs nobody ever sent and regenerates itself as the API evolves. Traffic-derived, if your spec is missing or fictional, because at least it tests what actually happens. The worst position is paying for neither and relying on whatever a teammate last hand-copied from the docs.

Pricing and plan details mentioned here are as published by the vendors around September 2026 and can change. Confirm on the official Schemathesis, Keploy, Postman, and Kiro sites before budgeting.

Pricing and plan details are as published by the vendor around September 2026 and can change - confirm on the official site.

Share this article

Related articles

Continue exploring similar guides and insights