Your Model Upgrade Is a Breaking Change: Build Contract Tests for LLM Providers in TypeScript
Most code that calls a model has one line that looks harmless.
model: "claude-sonnet-5"
Changing it feels like a config change.
But that string is part of an API contract.
And this month, the contracts changed.
* On September 17, Google released Antigravity Agent 09-2026. If you run tools locally or parse `function_call` steps, "the built-in tools changed": PascalCase parameters, and `write_file(path, content)` became `write_to_file` or `replace_file_content`. The old `antigravity-preview-05-2026` "shuts down on October 5, 2026."
* On September 22, Anthropic launched Claude Opus 5.5. Per the release notes, `thinking: {"type": "disabled"}` and `{"type": "enabled", ...}` "return a 400 error." So do `tool_choice` types `any` and `tool`.
* On September 28, Anthropic launched Claude Sonnet 5.5. The release notes say "Code written for Claude Sonnet 5 can break on Claude Sonnet 5.5 in five ways."
* On September 29, at DevDay, OpenAI released GPT-6.1 Sol. Its model page says "The `none` and `minimal` reasoning efforts are not supported."
Here are Sonnet 5.5's five, in the release notes' words:
1. "To turn off up-front thinking, send `thinking: {"type": "between_tools"}` instead of `"disabled"`, at `high` effort or below."
2. "Forced tool use (`tool_choice` types `any` and `tool`) returns a 400 error."
3. "Thinking blocks are tied to the model and the conversation."
4. "On the Claude API and Google Cloud, the earlier `computer_20251124` computer use tool isn't accepted."
5. "The advisor tool rejects Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5 as advisors."
The What's new page adds one that "alters the response shape without failing any request": text between tool calls comes back in `thinking` blocks.
A 400 is loud.
An empty progress message is quiet.
Different companies.
Same pattern.
**Changing a model string is a dependency upgrade. It deserves a test suite.**
So let's build one.
No API key. Both providers are mocks: their request rules follow the docs above, and their replies are made up.
## Table of Contents
1. What We Are Building
2. Project Setup
3. Step 1: A Neutral Request and Response
4. Step 2: Mock Two Model Versions
5. Step 3: Write the Contracts
6. Step 4: Diff Two Runs
7. Step 5: Check Known Breaks From the Docs
8. Step 6: Build the Gate and Run It
9. Where It Breaks Down
10. The Bigger Idea
## What We Are Building
One check reads the release notes. The other diffs two model versions.
## Project Setup
You will need Node.js 18 or newer.
mkdir model-upgrade-gate
cd model-upgrade-gate
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as `upgrade-gate.ts`.
## Step 1: A Neutral Request and Response
type Req = {
prompt: string;
maxTokens: number;
thinking?: { type: "adaptive" | "disabled" | "between_tools" };
toolChoice?: { type: "auto" | "none" | "any" | "tool" };
tools?: string[];
};
type Block =
| { type: "text"; text: string }
| { type: "thinking"; thinking: string }
| { type: "tool_use"; name: string; input: Record<string, unknown> };
type Stop = "end_turn" | "tool_use" | "max_tokens" | "refusal";
type Ok = { status: 200; stopReason: Stop; content: Block[] };
type Res = Ok | { status: 400; error: string };
type Provider = { model: string; send: (req: Req) => Res };
Your app's shape, not a vendor SDK.
**Your contracts should describe your app, not the provider.**
## Step 2: Mock Two Model Versions
// MOCK PROVIDER. No network, no API key. The two Sonnet 5.5 rejections follow
// Anthropic's docs (Sep 28), and the tool_choice error is quoted from them.
// The other error text and every reply are made up.
const reply = (stopReason: Stop, ...content: Block[]): Ok => ({ status: 200, stopReason, content });
function mockClaude(model: "claude-sonnet-5" | "claude-sonnet-5-5"): Provider {
const v55 = model === "claude-sonnet-5-5";
const send = (req: Req): Res => {
if (v55 && req.thinking?.type === "disabled") {
return { status: 400, error: 'invalid_request_error: use "between_tools"' };
}
if (v55 && ["any", "tool"].includes(req.toolChoice?.type ?? "auto")) {
return { status: 400, error: 'tool_choice: type "tool" and "any" are not supported for this model.' };
}
if (req.prompt.startsWith("[refuse]")) return reply("refusal");
if (req.maxTokens < 50) return reply("max_tokens", { type: "text", text: "Q3 revenue grew" });
if (req.tools?.includes("get_weather")) {
const note = "Checking the forecast first. Then I'll compare it with yesterday.";
const shown = req.thinking?.type === "between_tools" ? note : "";
const progress: Block = v55 ? { type: "thinking", thinking: shown } : { type: "text", text: note };
return reply("tool_use", progress, { type: "tool_use", name: "get_weather", input: { city: "Paris" } });
}
if (req.tools?.includes("classify_ticket")) {
return reply("tool_use", { type: "tool_use", name: "classify_ticket", input: { label: "billing" } });
}
return reply("end_turn", { type: "text", text: '{"total": 42.5, "currency": "USD"}' });
};
return { model, send };
}
Sonnet 5.5 rejects `disabled` thinking and forced tool use.
Its progress note also moves into a `thinking` block. At the default `display: "omitted"`, the docs say its text is empty. With `between_tools`, it comes back.
A mock that agrees with everything is just a very polite liar.
## Step 3: Write the Contracts
type Contract = { name: string; req: Req; check: (res: Ok) => string | null };
const weather: Req = { prompt: "Weather in Paris?", maxTokens: 500, tools: ["get_weather"] };
const contracts: Contract[] = [
{
name: "output schema",
req: { prompt: "Extract the invoice total as JSON", maxTokens: 500 },
check: ({ content: [first] }) => {
const data = JSON.parse(first?.type === "text" ? first.text : "null");
return typeof data?.total === "number" && typeof data?.currency === "string" ? null : "bad JSON shape";
},
},
{
name: "tool call format",
req: weather,
check: (res) => {
const call = res.content.find((b) => b.type === "tool_use");
return res.stopReason === "tool_use" && typeof call?.input.city === "string" ? null : "bad tool call";
},
},
{
name: "progress text between tools",
req: weather,
check: ({ content: [first] }) => {
const shown = first?.type === "text" ? first.text : first?.type === "thinking" ? first.thinking : "";
return shown ? null : `user sees nothing before the tool call (empty ${first?.type} block)`;
},
},
{
name: "forced tool use",
req: { prompt: "Classify this ticket", maxTokens: 200, tools: ["classify_ticket"], toolChoice: { type: "tool" } },
check: (res) => (res.stopReason === "tool_use" ? null : "no tool call"),
},
{
name: "thinking off (fast path)",
req: { prompt: "Summarize in one line", maxTokens: 200, thinking: { type: "disabled" } },
check: () => null, // a 200 is the whole contract
},
{
name: "token limit stop reason",
req: { prompt: "Write the full quarterly report", maxTokens: 20 },
check: (res) => (res.stopReason === "max_tokens" ? null : `got ${res.stopReason}`),
},
{
name: "refusal behavior",
req: { prompt: "[refuse] a request the model declines", maxTokens: 200 },
check: (res) => (res.stopReason === "refusal" && res.content.length === 0 ? null : "refusal not clean"),
},
];
Seven promises. Each check returns `null` or a reason.
The refusal contract follows the docs: a declined request returns HTTP 200 with `stop_reason: "refusal"`.
## Step 4: Diff Two Runs
type Result = { name: string; pass: boolean; detail: string };
function runSuite(provider: Provider): Result[] {
return contracts.map(({ name, req, check }) => {
const res = provider.send(req);
if (res.status !== 200) return { name, pass: false, detail: `${res.status} ${res.error}` };
try {
const failure = check(res);
return { name, pass: failure === null, detail: failure ?? "" };
} catch (err) {
return { name, pass: false, detail: `threw: ${(err as Error).message}` };
}
});
}
function contractDiff(current: Provider, candidate: Provider) {
const before = runSuite(current);
const after = runSuite(candidate);
const broke: string[] = [];
console.log(`\nContract diff (MOCK ${current.model} -> MOCK ${candidate.model})`);
after.forEach((a, i) => {
const status = before[i].pass && !a.pass ? "BROKE" : a.pass ? "same" : "FAIL";
if (status === "BROKE") broke.push(a.name);
console.log(` ${status.padEnd(6)} ${a.name.padEnd(28)} ${a.detail}`.trimEnd());
});
return broke;
}
Only one transition matters: **passed before, fails now.**
## Step 5: Check Known Breaks From the Docs
type KnownBreak = { model: string; param: string; bad: string[]; docs: string; source: string };
// From the vendors' docs, checked Oct 3, 2026.
const knownBreaks: KnownBreak[] = [
{ model: "claude-sonnet-5-5", param: "thinking.type", bad: ["disabled"], docs: 'send "between_tools" instead', source: "Claude notes, Sep 28" },
{ model: "claude-sonnet-5-5", param: "tool_choice.type", bad: ["any", "tool"], docs: "returns a 400 error", source: "Claude notes, Sep 28" },
{ model: "claude-opus-5-5", param: "thinking.type", bad: ["disabled", "enabled"], docs: "returns a 400 error", source: "Claude notes, Sep 22" },
{ model: "gpt-6.1-sol", param: "reasoning.effort", bad: ["none", "minimal"], docs: "not supported", source: "OpenAI model page" },
{ model: "antigravity-preview-09-2026", param: "tools", bad: ["write_file", "read_file", "list_files"], docs: "built-in tools changed", source: "Gemini changelog, Sep 17" },
];
const shutdowns: Record<string, string> = { "antigravity-preview-05-2026": "2026-10-05" };
type CallSite = { site: string; from: string; to: string; params: Record<string, string[]> };
function checkKnownBreaks(sites: CallSite[], today: string) {
let count = 0;
console.log("\nKnown breaks (from release notes)");
for (const s of sites) {
for (const rule of knownBreaks.filter((r) => r.model === s.to)) {
for (const value of (s.params[rule.param] ?? []).filter((v) => rule.bad.includes(v))) {
count++;
console.log(` BREAK ${s.site}: ${rule.param}=${value}: ${rule.docs} [${rule.source}]`);
}
}
const end = shutdowns[s.from];
const days = (Date.parse(end) - Date.parse(today)) / 86_400_000;
if (end) console.log(` DEADLINE ${s.site}: ${s.from} shuts down ${end} (${days} days)`);
}
if (count === 0) console.log(" no known breaks");
return count;
}
This is the deprecated-params check. Every row comes from a vendor's docs, with its date.
The `DEADLINE` line isn't a failure. It's a reason to hurry.
## Step 6: Build the Gate and Run It
// Illustrative call sites in a made-up app.
const S5 = "claude-sonnet-5", S55 = "claude-sonnet-5-5";
const callSites: CallSite[] = [
{ site: "invoice-extractor", from: S5, to: S55, params: {} },
{ site: "ticket-classifier", from: S5, to: S55, params: { "tool_choice.type": ["tool"] } },
{ site: "fast-summary", from: S5, to: S55, params: { "thinking.type": ["disabled"] } },
{ site: "code-agent", from: "gpt-6-sol", to: "gpt-6.1-sol", params: { "reasoning.effort": ["none"] } },
{ site: "file-agent", from: "antigravity-preview-05-2026", to: "antigravity-preview-09-2026", params: { tools: ["write_file"] } },
];
const today = "2026-10-03";
console.log(`Upgrade gate, ${today}`);
const breaks = checkKnownBreaks(callSites, today);
const broke = contractDiff(mockClaude(S5), mockClaude(S55));
const blocked = breaks > 0 || broke.length > 0;
console.log(blocked ? `\nGATE: BLOCKED (${breaks} known breaks, ${broke.length} contract regressions)` : "\nGATE: OPEN");
process.exitCode = blocked ? 1 : 0;
Run it:
npx tsx upgrade-gate.ts
Real output:
Upgrade gate, 2026-10-03
Known breaks (from release notes)
BREAK ticket-classifier: tool_choice.type=tool: returns a 400 error [Claude notes, Sep 28]
BREAK fast-summary: thinking.type=disabled: send "between_tools" instead [Claude notes, Sep 28]
BREAK code-agent: reasoning.effort=none: not supported [OpenAI model page]
BREAK file-agent: tools=write_file: built-in tools changed [Gemini changelog, Sep 17]
DEADLINE file-agent: antigravity-preview-05-2026 shuts down 2026-10-05 (2 days)
Contract diff (MOCK claude-sonnet-5 -> MOCK claude-sonnet-5-5)
same output schema
same tool call format
BROKE progress text between tools user sees nothing before the tool call (empty thinking block)
BROKE forced tool use 400 tool_choice: type "tool" and "any" are not supported for this model.
BROKE thinking off (fast path) 400 invalid_request_error: use "between_tools"
same token limit stop reason
same refusal behavior
GATE: BLOCKED (4 known breaks, 3 contract regressions)
Exit code 1. CI stops.
Look at `progress text between tools`. No 400. The user just stops seeing progress.
**The quiet break is the one a status code will never catch.**
The docs name the fixes: `between_tools`, `auto` plus strict tool use, `low` instead of `none`, and the new Antigravity tool names.
## Where It Breaks Down
### Mocks Drift
I copied the rules by hand. They also differ by platform: `computer_20251124` is rejected on the Claude API and Google Cloud, but Sonnet 5.5 still accepts it on Amazon Bedrock.
Run the contracts against the real API before trusting a green gate.
### Behavior Isn't a Contract
Anthropic says Sonnet 5.5's "effort levels are recalibrated." A schema check can't see that. Evals can.
### State Needs Real Conversations
Thinking blocks are tied to the model, the conversation and the account. On newer accounts, replaying one after editing history can return a 400. Single requests miss that.
## The Bigger Idea
We already treat libraries this way.
Pin the version. Read the changelog. Run the tests. Then upgrade.
Models get a string change and a hopeful deploy.
┌──────────────────────────────────────────────┐
│ Upgrade gate │
│ │
│ Release notes ──→ Known breaks ──┐ │
│ ↓ │
│ Current ──→ Contracts ──→ Diff ──→ Gate │
│ Candidate ──→ Contracts ──┘ │
└──────────────────────────────────────────────┘
The model provides capability.
The release notes provide warnings.
The contracts provide expectations.
The diff provides evidence.
The gate provides a decision.
Three vendors, four releases, twelve days. I think upgrade gates become as normal as lockfiles.
That part is prediction, not history.
**A new model is a new dependency. Ship it like one.**
## **Your code has a history. Helix makes it understandable.**
I'm building Helix so every change, including a model upgrade, comes with evidence: what changed, why, and what it touched.
**Connect your GitHub and see what your code knows.**
Explore Helix →