Home/Blog/Checkout demo
Demo · Repository decisions

The tests passed. The payment ran twice.

A checkout feature passed every test its coding agent wrote. Retrying the same request still created a second charge. A separate session with three stored team decisions produced code that handled the retry correctly.

A retry should not become a second purchase. Your team may already know that rule. The question is whether your coding agent has it available when it implements the endpoint.

We built Acme Checkout, a small fictional reference app, to make that gap visible. Two Haiku 4.5 sessions received the same feature prompt, starting code, and tools. Both implemented the endpoint and wrote passing tests. One session also received three team policies as Metatron decision files in the repository.

Then we submitted the same simulated $12.99 payment twice, with the same client request key.

Same request, retried onceWithout decisionsWith decisions
Charge IDs returnedch_1, ch_2ch_1, ch_1
Simulated charges created21
Simulated total$25.98$12.99
Agent’s own testsPassPass

The gateway is an offline, in-memory adapter. No real payments were made.

Watch the code and the retry

Watch Part 2 on YouTube · 4 minutes. See the complete prompt, recorded agent read events, unmodified agent patches, and checks rerun against those outputs.

The short task left room for a missing rule

The task asked for a payment endpoint: validate the input, call the supplied gateway, return the charge, and add meaningful tests. Both sessions received the same instructions to check for repository decisions before inspecting application code.

The shared prompt specified positive integer amounts, supported currencies, and invalid-input handling. The additional retry-key and error-format requirements lived in the decision files supplied to the context session.

Read the exact prompt used in both sessions
Before inspecting or editing application code, check whether context.md exists. If present, read it and the decision files it lists, and follow those repository policies. State which decisions you used.

Implement checkout.api.create_payment(request, gateway) for POST /payments using the supplied Request, Response and payment adapter.

The request body contains amount and currency. Return status 201 with the gateway's charge data for a valid payment, or status 400 for invalid input. Use the injected gateway and preserve the public signatures. Add meaningful behavior tests in tests/test_create_payment.py and run them along with the existing tests. Do not add dependencies or modify the gateway, HTTP types, existing tests, or repository decisions. Implement the source function as well as the tests.

Request.body is already parsed JSON. Accept only a dictionary body, a positive integer amount (not a boolean), and currency "usd" or "eur". Return 400 for malformed bodies or invalid value types, including list/dict currencies. Use the actual provided Request and Response fields.

Run tests with ./run-tests tests/test_create_payment.py or ./run-tests tests. The wrapper uses the prepared local Python environment. Do not edit the wrapper, access the network, install dependencies, or commit.

The baseline met its feature brief. Its fourteen new tests and the four existing gateway tests passed. Its implementation called:

charge = gateway.create_charge(amount, currency)

The gateway already supported retry keys, but its interface made them optional. The API’s requirement to use a client-owned key was a team policy. That requirement needed to reach the agent.

A decision carries the rule and the reason

The context session’s repository index linked to three files: client-owned idempotency, integer money, and a shared API error format. The first decision made the retry behavior explicit.

Repeating the same client request must reuse its charge, while two distinct client keys must permit two intentional identical payments.

The same file required the endpoint to reject missing or blank keys before calling the gateway and to pass the supplied key unchanged. Its rationale explained both sides: avoid charging twice on a retry, while allowing a customer to make an intentional second purchase.

That distinction matters. Deduplicating by payment amount and currency would collapse two legitimate purchases. Generating a new key for every request would allow the retry to create another charge.

Metatron keeps this kind of operating knowledge in readable repository files, where a team can review the rule and its rationale together.

The agent read the files and changed the behavior

The recorded tool trace shows successful reads of the index and all three decisions before the context session’s first code edit. Its resulting charge call included the client key:

charge = gateway.create_charge(amount, currency, idempotency_key)

The agent also wrote a regression test that calls the endpoint twice with the same request. These are the key assertions from that test:

first = create_payment(request, gateway)
second = create_payment(request, gateway)
assert first.body == second.body
assert len(gateway.charges) == 1

All twenty-six new tests passed, alongside the four existing tests. The new tests also failed against the original unimplemented endpoint. The policy had reached both the feature and the tests written for it.

The independent checks agreed

We evaluated both implementations with eleven policy checks kept outside the agents’ working directories. They covered valid input, malformed input, retry behavior, key handling, intentional repeat purchases, and error responses.

TrialWithout decisionsWith decisions
First paired run5/1111/11
Unchanged repeat5/1111/11

The second pair used the same prompt, policies, model, and checks. The duplicate-charge replay matched too: two simulated charges without the decisions, one with them. We made no manual repairs to either agent’s patch.

This is a selected reference demonstration, developed through exploratory trials. It shows an agent following additional supplied team requirements; it does not establish an average improvement across coding tasks or a unique advantage over every other way of providing those requirements.

Set up the decisions before the next task

Part 1 shows how those rules become repository files. It starts with the Acme Checkout fixture, runs metatron context setup, and shows the generated changes. The files-first workflow needs no Metatron server or MCP connection.

Watch Part 1 on YouTube · 3½ minutes. Setup, the full authoring prompt, decision diffs, and validation.

We then give the agent the three declared team rules and ask it to capture each rule with its rationale. The video follows the authoring step with the actual decision diffs and the updated index, then runs metatron files lint --path context/decisions to check the format.

That check validates the files’ structure. Human review still decides whether the rules are right. With the default pr review gate, decision files are authored on a working branch and become canonical through a reviewed pull request. The video stops with the files ready for that review; the application and tests remain unchanged.

The setup walkthrough was recorded after the comparison trials. It uses the same policy wording, organized under the documented Pattern and Rationale headings. These episodes show declared requirements becoming files and guiding a task; they do not show an agent learning a policy from an earlier failure.

Give the next session the decisions it needs

A passing test suite tells you the implementation meets the expectations those tests encode. In this demo, the context session had the team’s retry policy available, read it, implemented it, and tested it.

Think about a rule your reviewers repeatedly have to explain: which adapter a feature must use, how retries should behave, or why an apparently simpler approach is wrong for your system. Record the rule and its reason, then make that decision available before the next agent starts editing.