Salud Capital
Salud Capital · Research
September 15, 2026 · GREEN
AI Operations · Spend & Vendor Research

A Real-World Test: Running Claude Max at $200 Without Running Out, and Where Other Vendors Fit

How to run the $200 Max plan through the month at 2 to 10 active hours a day, which work to move to cheaper models, and where DeepSeek, Kimi, GLM, Qwen and the US hosts that serve them fit.

Claude MaxCapacity scenariosModel routingOpen weightsUS hostsPRC jurisdiction
Salud Capital Research · September 15, 2026 · A real-world test of one principal's own AI usage
Read this first · these dollar figures are not charges

Every "$" figure for Claude usage in this report is what that usage would cost if bought through the API. It is a unit for measuring usage, not money charged. The Max plan covers usage up to its limits for the $200 monthly fee. Money beyond the fee is charged only when usage credits are on and usage passes the limit, and the account's monthly spend cap stops those charges (it did on 2026-08-26). Actual charges are on claude.ai under Settings → Usage. The dollar unit is used because Anthropic publishes limits in no other measurable form.

About this test

The usage figures below were measured on one principal's own AI subscription over two weeks in September 2026 and are reported as recorded, with names and internal systems removed. Source discipline: every vendor price, plan term, government action and standard was read on the publisher's own page on 2026-09-15 before it was used. Estimates are labelled estimate and carry their formula. Secondary figures that could not be verified were dropped; the list is in §10.

01   Executive Summary

The answer in one page

The $200 Max plan is not a flat price for unlimited work. It is a metered pool with two ceilings: a five-hour session limit and a weekly limit, shared by every Claude surface.12 From two limit events on this account, those ceilings sit near $487 of API-equivalent usage per five-hour window and $2,899 per week — roughly $12,600 a month (estimate: $2,899 × 30.44 ÷ 7).68 Those are single observations of an algorithmic system, not published constants.

Claude Code usage since 9/1
$5,221
API-equivalent, 25,512 priced messages
Cache reads share of cost
72.6%
re-reading long conversations
Calibrated ceilings
$487 / $2,899
per 5 hours / per week
Modeled saving, all levers
46–53%
at 6 h/day × 5 days

Where the plan runs out today, and after optimization (15% of usage assumed to come from Claude chat, Design and desktop, which the transcripts do not capture; full sensitivity in §04):

Active hours per dayCurrent habits — 5-day weekCurrent — 7-day weekOptimized — 5-day weekOptimized — 7-day week
2Fits (52% of weekly)Fits (73%)Fits (25%)Fits (35%)
4Fits (80%), heavy 5-hour windows at 92%Runs out ~day 7 (112%)Fits (38%)Fits (53%)
6Runs out day 5 (108%); 5-hour limit hit on heavy daysRuns out day 5 (151%)Fits (51%)Fits (71%)
8Runs out day 4 (136%)Runs out day 4 (190%)Fits (64%)Fits (89%) — tight
10Runs out day 4 (164%)Runs out day 4 (230%)Fits (77%)Over by ~8% (108%, day 7)

Optimized = Sonnet 5 as the default with Opus used for about 15% of cost-weighted work, Haiku for background subagents, conversations reset or compacted well before they reach hundreds of thousands of tokens, and no more than two background agents at once. Numbers are from the scenario model in §04 and are estimates.

What to move where.

  1. Change the default model. On a Max plan Claude Code starts every session on Opus 5 unless told otherwise,13 and this account's measured mix was 66.6% Opus and 28.0% Fable by cost, with Sonnet at 5.3% and Haiku at 0.06%.68 Anthropic's own guidance is Sonnet for most coding and Opus when the problem is genuinely hard.712 Modeled saving: 21–30% of weekly usage.
  2. Stop re-reading enormous conversations. Opus on a Max plan runs with a one-million-token context and compacts only near ~967K tokens by default.13 The ten most expensive sessions were re-reading 315K–832K tokens per message by their final tenth.68 One Opus 5 turn at 550K tokens costs $0.275 in cache reads alone; at 150K on Sonnet 5 it costs $0.03 — about 9× less (estimate from published prices).9 Set a compaction window, clear between tasks.
  3. Put background agents on Haiku or Sonnet, and cap how many run at once. Claude Code allows 20 concurrent subagents by default, adjustable with one environment variable.14 Concurrency does not change weekly usage; it decides whether a burst trips the five-hour limit.
  4. Treat Fable as an explicit choice. Fable 5 costs twice Opus per token, its cache reads cost twice as much, its thinking cannot be switched off, and on Max it may use at most half the weekly limit.6913 It was 28% of measured cost.
  5. Keep the usage-credit monthly cap on. Past the limits, usage bills at standard API rates.4 At the 2026-09-08 to 09-15 pace, uncapped overflow is estimated at $3.2K, $6.0K or $10.0K a month depending on whether 0%, 15% or 30% of usage happens outside Claude Code — a floor that ignores five-hour overflows. The August 26 "monthly spend limit" notice shows a cap is already in force on this account.68
  6. Other vendors are an overflow and bulk-task lane, not a replacement. Open-weight models served by US hosts range from a few percent of Sonnet 5's per-token price (gpt-oss-120b, DeepSeek Flash-class) to more than Sonnet 5 (Kimi K3).39404446 Our own bench did not test any of them (no keys), so no production task moves until a second bench round runs. Financials, legal files, client documents, pricing and credentials never leave the current Anthropic arrangement.

Expected savings. At 6 active hours a day, 5 days a week, the combined levers take modeled weekly usage from $2,660 to $1,253–$1,429 (−46% to −53%), which moves the plan from "runs out on day 5" to "about half the weekly limit used." None of the levers costs money; each is a setting or a habit.

02   Observations

What the measurements show

2.1 Measured usage (Claude Code transcripts, 2026-09-01 to 2026-09-15)

The principal's own telemetry covers every Claude Code session and subagent file on the principal's workstation: 1,502 files, 25,512 priced assistant messages since 2026-09-01, priced at Anthropic's published per-token rates.689

MeasureValueWhat it means
API-equivalent usage$5,220.58What the same tokens would cost on the API; the plan meters against this kind of consumption
Cost by modelOpus 66.6% · Fable 28.0% · Sonnet 5.3% · Haiku 0.06%Opus is the default on Max,13 and nearly nothing runs on Haiku
Cache-read tokens as share of cost72.6%Long sessions re-send the whole conversation every turn; this outweighs model choice
Background agents (subagents)33.2% of costOne third of usage is agents the principal did not type to directly
Rolling 5-hour windowsmedian $91 · p95 $324 · max $486The max window is the one that hit the session limit on 9/8
Active human hours per day0.7 to 11.2Heaviest on 9/1 and 9/14
Largest single session$1,133, 1,926 messagesOpen across ~264 calendar hours
Week 37 (Sep 7–13)$2,123 for 20.5 active hours = $103.5 per active hourThe baseline week for the scenarios

Two corrections to the input notes. The notes cite "97,361 unique priced messages (25,512 since 9/1)"; the 97,361 figure is every unique assistant message in the scanned files with no date filter, and 25,512 is the priced count since 2026-09-01 (it equals the sum of the per-day counts). The notes give the background share as 32.8%; the data file gives 33.2%.

Coverage gap. These transcripts see Claude Code only. Claude chat, Claude Design, the desktop chat and anything else signed in with the same account draw on the same pool, and Anthropic says so explicitly: "your usage of all different Claude product surfaces … counts towards the same usage limit."2 Every number above is therefore a floor. the principal's description of the situation — the same usage pattern in a different chat window — matches Anthropic's.

2.2 The internal telemetry report

the firm's own telemetry for 2026-09-08 to 2026-09-15 found $4,158 of subscription usage at API rates in 8 days (37.8 active hours, about $110 per active hour) against ~$103 of real spend on the platform's API key.69 Its first pass reported $1,523 and missed 493 background-agent files; the corrected figure is the one used here.

Two items in that report need adjusting against the verified price sheet:

  • Opus 4.8 is priced at $5/$25 per million tokens, not $15/$75.9 the internal pricing table (internal pricing reference) still carries $15/$75.71 At the published price, cockpit chat on Opus 4.8 cost about $9.64, not $28.93, and the API-key total is about $84 (estimate: $28.93 × 5 ÷ 15). The report itself flagged this possibility.
  • Fable does not appear in the telemetry report, while the wider usage pass attributes $411 of Fable usage to 9/8 alone and 28% of all cost to Fable. The two passes classify models differently, and it is not established whether the Fable usage was Fable 5 or Fable 5.1 (whose cache reads cost a quarter as much).9 This is unreconciled and is listed in §10.
2.3 The near-miss extra-usage bill

The telemetry report warned that this week's pace, if billed at API rates, implies a bill of $10,000 or more a month, and the principal caught the usage warning. The mechanism is real: once usage credits are enabled, "subsequent usage will be billed at standard API pricing rates."4 The size depends on how much usage exceeds the limit, not on total usage:

Unmeasured share (chat, Design, desktop)Weekly usage vs calibrated weekly limitEstimated monthly overflow if uncapped
0%126%~$3,200
15%148%~$6,000
30%179%~$10,000

Estimate: weekly pace = $4,158 ÷ 8 × 7 = $3,638; plan-wide = pace ÷ (1 − unmeasured share); overflow = (plan-wide − $2,899) × 30.44 ÷ 7. This ignores five-hour overflows on bursty days, so it is a floor. It is also not the $15,600 "per month at this pace" figure, which prices all usage rather than only the part past the limit.

Two safeguards already exist in Anthropic's product. A monthly spending cap can be set on usage credits (or set to unlimited),4 and the account hit that cap on 2026-08-26 about $158 past the weekly limit.68 Discounted usage bundles of $50, $250 and $1,000 carry 10%, 20% and 30% discounts, up to $2,000 a month on individual plans.5

2.4 What the bench showed about sufficiency per task type

The model bench ran eight graded suites across Anthropic and OpenAI models and effort levels for $2.23.70 "Sufficient" means the cheapest configuration within 95% of the best score (100% for the two engineering suites).

SuiteBest configuration (pass)Cheapest sufficient configurationOpus 5 best effort, $/case ÷ sufficient $/caseHow to read it
aisc (steel shape lookups)Opus 5 low (94.4%)Opus 5 low1.0×Deterministic data — use a lookup table in code, not a model
units (unit conversion)Haiku 4.5 (100%)Haiku 4.526×Deterministic — code first
classifyGPT-5 mini low, Haiku 4.5 and others (100%)GPT-5 mini low19×Small models suffice
memory-classifySonnet 5 low/medium (60%)Sonnet 52.7×Best is only 60%: a labelling or task-definition problem, not evidence any model is fine
oversightSonnet 5 and Opus 5 (100%)Sonnet 5 low2.5×Constructed cases, n=18
grounded-qaGPT-5 mini low (94.1%)GPT-5 mini low12×Constructed cases, n=17
refuse-vs-guessSeveral (100%)GPT-5 mini low12×Constructed cases, n=18
code-fixGPT-5 mini low, Haiku, Sonnet (100%)GPT-5 mini low11×n=8; Opus 5 scored 62.5–75%

What the bench supports, read carefully:

  • Opus 5 was the sole top scorer on one suite (aisc), tied cheaper configurations on five, and was beaten on two (memory-classify, code-fix). On units and code-fix its score fell as effort rose (units 100% → 94% → 82%; code-fix 75% → 62.5% → 62.5%). With 8–20 cases per suite, a one-case difference is noise; the direction, not the decimal, is the finding.
  • Oversight, grounded-QA and refuse-vs-guess cases were written from rules, and code-fix has eight cases. They are directional evidence, not production proof.
  • Coverage gaps, stated plainly: xAI did not run (the team had no API credit). Google ran partially: two of the models named in the harness were retired for new users and the connected key sat on free-tier rate limits, so Gemini results rest on 1–8 cases per suite and are not used here. DeepSeek, Qwen, Kimi, GLM, MiniMax, Doubao and OpenRouter were not run because no keys exist. Published benchmarks are not substituted for these gaps.
  • Google's free tier is not a private test bed. Google's price sheet marks free-tier content "Used to improve our products: Yes."21 The bench cases were synthetic, which limits the exposure, but no business data should ever touch a free-tier key.

An operational gotcha worth a paragraph. gpt-5 at high reasoning effort scored 0% on aisc and 5.6% on refuse-vs-guess. The model was not wrong; the calls ended with empty answers, because at a 300-token output ceiling its hidden reasoning consumed the entire budget before any visible text was written (the harness recorded output tokens at the cap with an empty answer string). GPT-5 mini at high effort showed the same failure less severely. Anthropic's adaptive-thinking models did not show it at the same ceilings. The practical rule: raising reasoning effort on a short-answer task can make results strictly worse unless the output limit rises with it, and raising the limit raises worst-case cost. Claude Code's own documentation makes the parallel point that thinking tokens are billed as output tokens and that lower effort is the cost lever for simpler tasks.12

03   The Claude Max Plan

Limit structure, shared pool, extra usage

3.1 What Anthropic publishes
TermAnthropic's published position (fetched 2026-09-15)
PriceMax 5x $100 a month; Max 20x $200 a month; monthly billing only1
Multiple"Max 20x provides 20 times more usage per session than the Pro plan"1
Five-hour limit"Your session-based usage limit will reset every five hours"1
Weekly limitA weekly limit across all models that "resets at a fixed time each week that is assigned to your account"1
Discretionary limitsAnthropic "may limit your usage in other ways, such as weekly and monthly caps or model and feature usage"1
Shared poolclaude.ai, Claude Code and Claude Desktop "count towards the same usage limit"; IDE usage too23
What drives usageConversation length and complexity, features, model and effort level2
Fable modelsIncluded on Max; "up to 50% of your weekly usage limits on Fable models," drawn from the same weekly limit6
Agent SDK and claude -pA June 2026 plan to move these to a separate monthly credit was paused; they still draw from subscription limits8
At the limitWait, upgrade, or buy usage credits23
Usage creditsBilled at standard API rates; monthly spend limit or unlimited; optional auto-reload; $2,000 daily redemption limit4
Usage bundlesPrepay $50 / $250 / $1,000 at 10% / 20% / 30% off; up to $2,000 a month on Pro and Max5
API key overrideAn ANTHROPIC_API_KEY environment variable makes Claude Code bill the API instead of the plan3
Terms for automation"Advertised usage limits for Pro and Max plans assume ordinary, individual usage of Claude Code and the Agent SDK"; plan credentials may not be used to route requests on behalf of other users10

No published page gives a number for either limit. The figures in this report come from this account's own limit events.

3.2 This account's calibrated limits
Event (Central time)What firedUsage at that moment (API-equivalent)
2026-08-25 08:39"hit your weekly limit"Week-to-date $2,899 (week boundary Wednesday ~9 pm CT, inferred from the notice)
2026-08-26 16:37"hit your monthly spend limit"Week-to-date ~$3,057: about $158 of usage credits past the weekly limit before the credit cap stopped it
2026-09-08 17:26"hit your session limit"Trailing five hours $481–494, matching the independently computed $486.48 peak window that hour

Source: The principal's own telemetry (limit-event scan of Claude Code transcripts; 16 account-limit notices and 5 subagent-concurrency notices after excluding false positives).68

Two observations on one account are the best available estimate, not guarantees. Anthropic describes limits as depending on conversation length, features, model and effort, and reserves the right to cap usage in other ways.12 A useful way to hold the numbers: the $200 plan bought roughly 63 times its price in API-equivalent usage at the weekly ceiling (estimate: $12,607 ÷ $200). The plan is excellent value inside its limits and bills at full API rates outside them.

3.3 Model defaults that shape usage on Max

These come from Claude Code's configuration documentation and explain much of the measured mix:

  • Default model on Max is Opus 5. On Pro it is Sonnet 5.13
  • Opus is automatically upgraded to a one-million-token context on Max, with standard pricing and no premium for tokens beyond 200K.13
  • Default compaction for 1M-context models happens near 967K tokens. The window is adjustable from 100K to 1M with /autocompact, --autocompact or CLAUDE_CODE_AUTO_COMPACT_WINDOW; CLAUDE_CODE_DISABLE_1M_CONTEXT=1 holds sessions to 200K.13
  • Default effort is high on every model that supports effort; maxEffortLevel caps it.13
  • Built-in Explore subagents inherit the main conversation's model (capped at Opus); custom subagents take a model: field or CLAUDE_CODE_SUBAGENT_MODEL.14
  • Twenty subagents may run concurrently by default; CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS changes it.14
  • The status line receives rate_limits.five_hour.used_percentage and rate_limits.seven_day.used_percentage, and /usage attributes plan usage to subagents, skills, plugins and MCP servers.1512

Anthropic's help article for Claude Code states "Sonnet is the default" while the Claude Code configuration page lists Opus 5 as the Max default.713 The configuration page is the more specific source and matches the measured mix.

04   Capacity Scenarios

2 to 10 active hours a day, 5 and 7 days a week

4.1 How the model works

The scenario model scales the baseline week (Week 37: 20.5 active hours, 6 active days) linearly:68

  • weekly usage = $69.09 × active hours (main sessions) + $117.57 × active days (background agents);
  • plan-wide usage = Claude Code usage ÷ (1 − unmeasured share), with the unmeasured share set to 0%, 15% or 30% by assumption;
  • the heavy five-hour window = the observed p95 window ($324), scaled by target hours per day ÷ the week's average, with the same model-mix multiplier;
  • the week "runs out" on the first day on which cumulative usage passes $2,899.

Optimized applies: Opus reduced to 15% (or 30%) of cost-weighted main-session work with the rest at Sonnet prices (Sonnet costs 1/2.5 of Opus per token); context resets that remove 33.7% of main-session cache-read cost; Haiku for background agents (1/2 of Sonnet); and, for the five-hour figure only, background concurrency capped at 2 instead of the observed 20.914

4.2 Weekly usage as a share of the calibrated weekly limit (15% unmeasured)
Hours/day × daysCurrent mixRuns outOptimized, Opus 30%Optimized, Opus 15%Runs out (optimized 15%)
2 × 552%27%25%
2 × 773%38%35%
4 × 580%43%38%
4 × 7112%day 760%53%
6 × 5108%day 558%51%
6 × 7151%day 581%71%
8 × 5136%day 473%64%
8 × 7190%day 4103% (day 7)89%
10 × 5164%day 489%77%
10 × 7230%day 4124% (day 6)108%day 7
4.3 Heaviest five-hour window as a share of the $487 limit (15% unmeasured)
Hours per dayCurrent mixOptimized, Opus 30%Optimized, Opus 15%
246% ($223)18%15%
492% ($446)35%30%
6137% ($669)53%45%
8183% ($892)70%60%
10229% ($1,115)88%75%

The five-hour limit depends on hours per day, not days per week, so 5-day and 7-day rows are identical. At current habits, any day above about 4 active hours with parallel agents risks the session limit. After optimization, even a 10-hour day stays under it in the model, but only because background concurrency is capped.

4.4 Sensitivity to the unmeasured share
Hours/day × daysCurrent: 0% / 15% / 30% unmeasuredOptimized (Opus 15%): 0% / 15% / 30%
2 × 762% / 73% / 88%30% / 35% / 42%
4 × 568% / 80% / 97%32% / 38% / 46%
4 × 795% / 112% / 136%45% / 53% / 64%
6 × 592% / 108% / 131%43% / 51% / 62%
6 × 7128% / 151% / 184%61% / 71% / 86%
8 × 5116% / 136% / 165%54% / 64% / 78%
8 × 7162% / 190% / 231%76% / 89% / 109%
10 × 5139% / 164% / 199%65% / 77% / 93%
10 × 7195% / 230% / 279%91% / 108% / 131%

Reading the grid. With current habits, the plan comfortably supports about 2 hours a day, and 4 hours a day on a 5-day week. With the optimized habits, it supports up to 10 hours a day on a 5-day week and up to 6–8 hours a day on a 7-day week, depending on how much usage happens outside Claude Code. Ten hours a day, seven days a week exceeds the weekly limit even when optimized — by about $219 a week at 15% unmeasured (estimate: $3,118 − $2,899), or roughly $950 a month of overflow at API rates before any bundle discount.

4.5 What the model does not capture

The model assumes weekly usage scales linearly with hours; that a smaller model finishes the same work in the same number of tokens (it may need more turns); and that context resets save a fixed fraction derived from the top-ten sessions' growth curves rather than a replay. It prices all non-Opus work at Sonnet rates, including the 28% that was Fable — which makes the savings from moving Fable work understated. The five-hour projections scale the shape of observed days. All of these are listed again in §10.

05   Levers

Ranked by estimated savings

Baseline: 6 active hours × 5 days, current mix, modeled at $2,660 a week ($2,073 main sessions, $588 background agents). Savings are sequential where noted, so they do not simply add.68

RankLeverHow (verified setting or habit)Modeled weekly savingConfidence
1Sonnet 5 as default; Opus for judgment/model sonnet or ANTHROPIC_DEFAULT_MODEL; opusplan plans with Opus and executes with Sonnet137$570 (Opus → 30%) to $803 (Opus → 15%) = 21–30%Medium: assumes equal token volume
2Compact early, clear between tasks/autocompact 200k (100K–1M allowed); CLAUDE_CODE_AUTO_COMPACT_WINDOW for launched agents; /clear per task; CLAUDE_CODE_DISABLE_1M_CONTEXT=1 where long context is not needed137$507 alone (19%); $310–367 after lever 1Low–medium: rough estimate, likely understated given 72.6% cache-read share
3Haiku (or Sonnet) for background agentsmodel: haiku in subagent frontmatter; CLAUDE_CODE_SUBAGENT_MODEL; Explore otherwise inherits the session model1412$294 (11%)Medium for exploration and search; not for judgment tasks
4Fable by explicit choice onlyFable is never the account default, but a /model choice persists to later sessions13Not in the scenario model. Moving the period's $1,464 of Fable usage to Opus saves ~$732 over the two weeks (estimate: × (1 − 5/10)); to Sonnet ~$1,171Medium on arithmetic; quality impact unknown
5Cap concurrencyCLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS (default 20); a cap in the agent launcher for launched agents14$0 weekly; keeps heavy days under the $487 five-hour limit (137% → 45–53% with levers 1–3)Medium
6Lower effort on routine work/effort; maxEffortLevel managed setting; default is high13Not quantified; the bench showed no accuracy gain from higher effort on its task typesLow (small bench)
7Schedule across windowsStart agent batches right after a five-hour reset; spread the same weekly hours over 7 days instead of 5$0 weekly; roughly halves the heaviest daily concentrationMedium
8Usage monitor with alertsStatus line reading rate_limits.five_hour and seven_day percentages; /usage attribution by subagent and MCP server; alert at 70% and 90%1512Prevents overflow rather than reducing usageHigh (feature is documented)
9Overflow routingKeep the usage-credit monthly cap; bundles at 10–30% off for planned overflow; the platform's own jobs on its API key with Batch at 50% off; bounded public-data jobs to US-hosted open weights after a bench run459Up to 30% off overflow dollars (bundles), 50% off batchable API workMedium

Combined: levers 1 (Opus → 30%), 2 and 3 take the week to $1,429 (−46%); with Opus at 15% the week is $1,253 (−53%). The input notes labelled the −46% combination as using the 15% Opus lever; recomputation shows it used the 30% lever (§10).

Why lever 2 matters more than it looks. Every turn re-sends the conversation. The per-turn cache-read cost at published prices:9

Context carried per turnOpus 5 ($0.50/M)Sonnet 5 ($0.20/M)Haiku 4.5 ($0.10/M)Fable 5 ($1.00/M)
150K tokens$0.075$0.030$0.015$0.150
550K tokens$0.275$0.110$0.055$0.550
967K tokens (default compaction point)$0.483$0.193$0.097$0.967

A 1,900-message session at 550K tokens of Opus context carries roughly $520 of cache reads (estimate: 1,900 × $0.275) before any output is written. Anthropic's guidance calls /clear "the single most effective lever for both quality and cost."7

06   Other Vendors

Plans, APIs, Chinese labs, US hosts, jurisdiction

6.1 Consumer and coding plans
VendorPlans and monthly price (verified)Coding agent on the planAutomation terms
AnthropicMax 5x $100 · Max 20x $2001Claude Code, shared limits3Ordinary individual use; no automated access except via API key or explicit permission1011
OpenAIPlus $20; Pro $100 and Pro $200 (new sign-ups and upgrades to Pro $200 paused from 2026-09-10)19Codex on all ChatGPT plans; one allowance shared with ChatGPT Work and other agent features20Not reviewed in this pass (policy pages returned 403)
GoogleAI Plus $4.99 · AI Pro $19.99 · AI Ultra $99.99 (5× Pro limits) and $199.99 (20×)2223Gemini CLI: 1,000 requests/day (Code Assist individual), 1,500 (AI Pro), 2,000 (AI Ultra); API-key free tier 250/day24Not reviewed in this pass
xAISuperGrok $30 · SuperGrok Plus $100; Heavy listed without a readable price26Grok Build, early beta for SuperGrok and X Premium Plus27Acceptable Use Policy prohibits "unauthorized automated or non-human means"28
Z.ai (GLM)GLM Coding Plan "starting at just 18 USD per month," with a five-hour and a weekly limit; works with Claude Code34Via Claude Code, Cline, OpenCodeSingapore-based controller (see 6.5)35

Where the fetched pages describe limits, they are metered allowances rather than unlimited use: five-hour and weekly limits at Anthropic and Z.ai, daily request quotas for Gemini CLI, and a shared allowance with paid credits for Codex. A second subscription would add a second pool to manage, not remove the need to manage the first.

6.2 Frontier API prices ($ per million tokens, input / output)
ModelInputOutputCached inputNotes
Claude Opus 55.0025.000.50Opus 4.5–4.8 same price; Batch 50% off; US-only inference 1.1×9
Claude Sonnet 52.0010.000.20Introductory price made standard; planned rise to $3/$15 cancelled9
Claude Haiku 4.51.005.000.109
Claude Fable 5 / Fable 5.110.0050.001.00 / 0.25Newer tokenizer on 4.7+ models yields ~30% more tokens for the same text9
GPT-6 Astra10.0050.001.00Short context; Batch/Flex half16
GPT-5.6 Sol4.0020.000.40Promotional pricing at least through 2026-11-2116
GPT-5.6 Terra2.0012.000.2016
GPT-5.6 Luna0.201.200.0216
GPT-5 / GPT-5 mini1.25 / 0.2510.00 / 2.00Models used in the bench1718
Gemini 3.1 Pro Preview2.0012.000.20Prompts ≤200K tokens21
Gemini 3.8 Flash0.753.750.075Through 2026-12-31; $1.50 / $7.50 from 2027-01-0121
Gemini 3.5 Flash1.509.000.15Model used in the bench21
Grok 4.62.006.00500K context25
6.3 Chinese labs, first-party APIs ($ per million tokens)
ModelVendorInput (cache miss)OutputCache hitNotes
DeepSeek V4.1 Flash (deepseek-flash)DeepSeek0.30 peak / 0.15 off-peak1.20 / 0.600.006 / 0.003Peak 01:00–04:00 and 06:00–10:00 UTC weekdays; 1M context; Anthropic-format endpoint29
DeepSeek V4 ProDeepSeek1.32 / 0.663.96 / 1.980.044 / 0.022Service continued past 2026-09-1429
Kimi K3Moonshot AI3.0015.000.30~1M context32
Kimi K2.7 CodeMoonshot AI0.954.000.1932
Kimi K2.6Moonshot AI0.954.000.1632
GLM-5.3Z.ai (Zhipu)1.404.400.2633
GLM-5.3-FlashZ.ai0.150.500.03GLM-4.7-Flash listed free33
Qwen3.8-MaxAlibaba Cloud Model Studio2.006.00Singapore deployment, "International" scope; the page lists Global, US and Chinese-mainland scopes for other models36
MiniMax-M3MiniMax0.301.200.06After a listed "permanent 50% off" ($0.60 / $2.40 list), ≤512K input37
Doubao / SeedByteDance (BytePlus ModelArk)Prices not readable in the fetched pages; not used38

Against Sonnet 5 ($2 / $10), DeepSeek V4.1 Flash at peak is 15% of the input price and 12% of the output price; GLM-5.3 is 70% and 44%; Kimi K3 is above Sonnet 5 on both. The large savings are in the Flash-class models, not the Chinese flagships.

6.4 The same open weights through US hosts ($ per million tokens, input / output)

Figures come from each host's own price page where readable, and otherwise from OpenRouter's public endpoints API, which lists the price and quantization of every upstream serving a model on OpenRouter.40

ModelFireworks44DeepInfra46Other listed routes (OpenRouter endpoints API)40
DeepSeek V4 Pro1.32 / 3.961.30 / 2.60Azure (US) 1.91 / 3.83 · Baseten 1.74 / 3.48 (fp4) · lowest listed: StreamLake 0.95 / 1.89 and Baidu 0.95 / 1.90 (fp8)
DeepSeek V4.1 Flash0.22 / 0.660.20 / 0.60 (fp8, via OpenRouter)Together 0.30 / 1.20 · DeepSeek first-party 0.15 / 0.60
Kimi K2.60.95 / 4.000.75 / 3.50Crusoe 0.70 / 3.50 (bf16) · lowest: Inceptron 0.46 / 2.61 (int4)
Kimi K33.00 / 15.00; "K3 US" 3.30 / 16.502.85 / 14.25Together 3.00 / 15.00 · Alibaba 3.45 / 17.25
GLM-5.31.40 / 4.400.90 / 3.00 (fp4, via OpenRouter)Together 1.40 / 4.40 · Z.AI 1.40 / 4.40
gpt-oss-120b0.037 / 0.17 (bf16, via OpenRouter)Groq, Together, Amazon Bedrock 0.15 / 0.60 · Cerebras 0.35 / 0.75
MiniMax-M30.28 / 1.10 (fp8, via OpenRouter)Together 0.30 / 1.20 · MiniMax 0.30 / 1.20

Hyperscalers also carry these weights: Amazon Bedrock lists DeepSeek (R1, V3.1, V3.2), MiniMax and Qwen models48; Microsoft lists DeepSeek-V4-Pro and DeepSeek-V4-Flash among models sold directly by Azure49; Vercel AI Gateway, on which Salud's apps already deploy, passes provider list prices through "no markup and no platform fee"50 (DeepSeek V4 Flash from $0.06 / $0.18).

Three observations from the endpoint data:

  1. The cheapest route is often a lower-precision one. OpenRouter labels quantization per endpoint (fp4, int4, fp8, bf16, or unknown). The input notes said precision is rarely disclosed; for OpenRouter's catalogue it usually is, although many endpoints still say "unknown."
  2. The cheapest route is often not a US company. The lowest-priced DeepSeek V4 Pro endpoints on 2026-09-15 included Baidu and Alibaba; Kimi and GLM routes include Moonshot AI and Z.AI directly. A router will use them unless told not to (6.5).
  3. The spread for one model is wide. DeepSeek V4 Pro output ranges from $1.89 to $4.90 across listed endpoints, and Kimi K3 output from $10.95 to $22.50 (fast variants). "DeepSeek is cheap" is a statement about a route, not a model.
6.5 Quality evidence
  • NIST's Center for AI Standards and Innovation (CAISI), May 2026: DeepSeek V4 Pro "lag[s] behind the frontier by about 8 months," is "the most capable PRC AI model evaluated by CAISI to date," performs on CAISI's evaluations "similarly to GPT-5," scores better on DeepSeek's self-reported evaluations than on CAISI's, and was more cost-efficient than GPT-5.4 mini on 5 of 7 benchmarks.56
  • CAISI, September 2025 (earlier DeepSeek models R1, R1-0528, V3.1): agents built on DeepSeek's most secure model were on average 12 times more likely to follow hijacking instructions than US frontier models; the model answered 94% of overtly malicious requests under a common jailbreak versus 8% for US reference models; and it echoed four times as many inaccurate CCP narratives.57 These findings concern older models and are cited as the government's most recent published security evaluation of the family, not as a measurement of V4.
  • Our own bench tested none of these models. No production task should move on published benchmarks alone.
6.6 Data jurisdiction and compliance

First-party Chinese APIs. DeepSeek's privacy policy names Hangzhou DeepSeek Artificial Intelligence Co., Ltd. as controller, states that it collects, processes and stores personal data in the People's Republic of China, and lists training and improving its models among the uses (with an opt-out right).30 Its Open Platform terms are governed by the laws of mainland China.31 Z.ai's privacy policy names a Singapore company as controller and says data is "generally processed in Singapore."35 The input notes' claim that Z.ai excludes API content from training by default was not found in that policy and is not repeated here.

PRC law, from the National People's Congress texts. National Intelligence Law Article 7 (2017 text): all organizations and citizens shall support, assist and cooperate with national intelligence work in accordance with law, and keep its secrets (our translation).66 Data Security Law Article 35: when public security or state security organs lawfully obtain data for national security or criminal investigation, the organizations and individuals concerned shall cooperate (our translation).65 The Personal Information Protection Law, in force since 2021-11-01, sets conditions for sending personal information outside China (Article 38).64 Authoritative text is Chinese; the NPC hosts an English PIPL page.67 This is a description of statutes, not legal advice; counsel should read it before any decision that depends on it.

Routers. OpenRouter can exclude providers that may store data (data_collection: "deny"), restrict routing to zero-data-retention endpoints (zdr), and allow or ignore named providers per request; an account setting excludes providers that train on prompts; enterprise accounts can pin processing to US or EU regions.4142 Its documentation also says it "does not have routing rules that change based on data retention policies of providers" — the zero-data-retention flag is the control, not a retention-window filter as the input notes described.42 Per-key credit limits with a configurable reset cap spend on a key; an exhausted limit returns an error rather than a bill.43

Hosts. Fireworks states it "does not log or store prompt or generation data for open models, without explicit user opt-in."45 DeepInfra's privacy notice states its security measures "comply with SOC 2 and ISO 27001 standards."47 Together AI, Groq and Baseten publish trust centers whose content did not render for our fetcher.515253 No SOC 2 report or ISO certificate was inspected for any vendor in this report; a trust-center page is a claim, and the report behind it is the evidence. Novita's jurisdiction (a Singapore entity with a San Francisco address, per the input notes) is unresolved; it appears as an upstream on OpenRouter routes above and should be excluded until resolved.

Anthropic. US-only inference on the API is available at a 1.1× price multiplier.9

6.7 Government actions and standards
BodyInstrumentRelevance
U.S. House of RepresentativesH.R. 1121, "No DeepSeek on Government Devices Act," introduced in the 119th Congress and referred to the Committee on Oversight and Government Reform58Signals federal posture; later status not verified (congress.gov returned 403)
State of New YorkGovernor Hochul's 2025-02-10 ban on DeepSeek on ITS-managed government devices and networks59State-level precedent relevant to public-sector clients
NIST / CAISIDeepSeek V4 Pro evaluation (May 2026) and DeepSeek models evaluation (Sept 2025)5657Capability and security evidence
NISTAI Risk Management Framework 1.0, under revision per the White House AI Action Plan54Voluntary governance frame for vendor selection
NISTAI 600-1, Generative AI Profile (2024-07-26)55GenAI-specific risk actions
ISO/IECISO/IEC 42001, AI management systems60Certification to ask vendors about (page returned 403; link only)
AICPASOC suite, including SOC 261The report to request, not the badge
GSA FedRAMPFedRAMP and its Marketplace62Relevant only if work touches federal systems
European UnionRegulation (EU) 2024/1689, the AI Act, of 13 June 202463Relevant to EU-routed processing
PRC National People's CongressPIPL, Data Security Law, National Intelligence Law646566Legal basis of the jurisdiction concern
07   Routing Policy

Which work goes to which tier

TierWhereTask classesGuardrails
0 — Code, no modelDeterministic functions and tablesdomain reference shape weights and properties, unit conversions, arithmetic roll-upsThe bench's aisc and units suites are deterministic; a model is the wrong tool
1 — Claude Max plan, interactiveClaude Code, chat, Design on the principal's own loginArchitecture, cross-cutting changes, hard debugging, review of drafts, anything needing judgment; all never-offload work belowSonnet 5 default; Opus by choice; Fable by explicit choice; compaction window set; monitor at 70% / 90%
2 — Claude API on the platform's keyinternal agent services with a Console spend limitCockpit chat (Sonnet 5), memory classifier (Haiku 4.5), oversight (Sonnet 5 low), scheduled research, batchable jobs via the Batch API at 50% off9Separate key per service; no ANTHROPIC_API_KEY on the principal's interactive shell, where it would silently switch Claude Code to API billing3
3 — US-hosted open weightsFireworks or DeepInfra direct, or OpenRouter with ignore on non-US and unresolved providers, data_collection: "deny", zdr: true, and a per-key credit limit4143Public or non-confidential bulk work only: tagging public catalogue data, summarizing public product literature, drafting public marketing copy, chores on open-source codeOnly after the bench's second round clears the task type; nothing from the list below
Never offloadStays on the current Anthropic arrangement; never to a first-party Chinese API, a free tier, or an unvetted hostFinancial records · legal and counsel files · client documents and client documents · pricing, proposals and costs · credentials, keys, .env values · patent and trade-secret materialExisting internal policies already restrict where these travel; this list adds no exceptions
08   Recommendations

Numbered, with owner and cost

#RecommendationOwnerCostExpected effect
1Set Claude Code's default model to Sonnet 5; use opusplan or pick Opus deliberately for design and hard debugging13Principal (user settings); platform (launcher defaults)$0−21% to −30% of weekly usage (modeled)
2Set /autocompact 200k (or 300K) in user settings and CLAUDE_CODE_AUTO_COMPACT_WINDOW for launched agents; /clear between unrelated tasks13Principal; platform$0−12% to −19% further (modeled, rough)
3Give every custom subagent an explicit model:haiku for search and exploration, sonnet for building; set CLAUDE_CODE_SUBAGENT_MODEL in launched sessions14Platform$0−11% (modeled)
4Set CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS=4 normally and 2 on heavy days; add the same cap to the cockpit launcher14Platform$0Keeps heavy days under the five-hour limit
5Make Fable opt-in per task; review why it carried 28% of usage6Principal$0Up to ~14% of total usage if that work moves to Opus (estimate)
6Build a status-line and internal alert on rate_limits.five_hour and seven_day at 70% and 90%; add plan-usage percentages to the weekly telemetry report15PlatformA few build hoursEarly warning before overflow
7Keep the usage-credit monthly cap at a figure the principal chooses; buy bundles only for planned overflow (10–30% off)45PrincipalCapped by choiceBounds the extra-usage bill
8Fix the internal pricing table: Opus 4.8 at $5/$25, add Fable 5/5.1 and Sonnet 5 at published prices; reconcile Fable attribution between the telemetry report and the usage pass971Platform$0Correct spend reporting (the internal agent platform key ~$84, not $103, for 9/8–9/15)
9Move cockpit chat off Opus 4.8 to Sonnet 5; keep the classifier on Haiku 4.5; run oversight on Sonnet 5 low (the bench's sufficient configuration)70Platform$0Lower API-key spend
10Replace model calls for domain reference lookups and unit conversions with deterministic codeEstimatingBuild hoursRemoves a 94%-accurate model from a 100%-knowable task
11Fix memory-classify labels or task definition before choosing its modelPlatformBuild hoursA valid bench for the classifier
12Bench round two under the same $35 cap: add OpenRouter (PRC-company and unresolved providers ignored, zdr on), a paid Google key and xAI credit, plus one agentic multi-step suitethe platform team runs; the principal approves and creates any accounts≤$35 plus prepaid creditsEvidence for Tier 3 before any production move
13If OpenRouter is adopted: per-key credit limit with a reset period, data_collection: "deny", zdr: true, provider ignore list, no business data4143Principal / security$0No surprise bill; no unintended jurisdiction
14Do not add ChatGPT, Google AI or xAI subscriptions now; revisit only if the optimized habits still exceed the weekly limit at the hours the principal choosesPrincipal$0 nowAvoids a second pool to manage
15No first-party Chinese API for any business data; counsel review before any exceptionPrincipal / counselCounsel time if raisedKeeps the never-offload list intact
16If many concurrent agents remain the primary workflow, ask Anthropic whether Max remains the right arrangement or whether Team Premium, Enterprise or API billing fits better, since advertised Max limits assume ordinary individual use10PrincipalA callRemoves ambiguity on plan fit
09   Conclusions

What this adds up to

  1. The plan is not the problem; the defaults are. Opus 5 as the Max default, a one-million-token window that compacts near 967K, high effort and twenty concurrent subagents are each reasonable alone. Together they turn a $200 plan with roughly $12,600 of monthly capacity into one that runs out on day 4 or 5 at 6–10 active hours a day.
  2. The biggest line item is re-reading. Cache reads are 72.6% of usage. Shorter conversations and a cheaper default model address the same cost from two sides, and neither needs a new vendor.
  3. With the levers applied, the plan fits nearly every working pattern modeled. Up to 10 hours a day on a 5-day week and 6–8 hours on a 7-day week stay under the calibrated limits; only 10 hours a day, 7 days a week needs paid overflow, estimated near $950 a month before bundle discounts.
  4. The bench is a good start and a small one. It shows cheap configurations matching or beating Opus 5 on bounded tasks and exposes a real reasoning-budget failure mode, but its constructed cases, small samples and missing vendors mean it cannot yet justify moving production work to other vendors.
  5. Cheaper models are real, and so are their costs. Flash-class open-weight models via US hosts are an order of magnitude cheaper per token than Sonnet 5. The trade is capability (CAISI puts the strongest PRC model about eight months behind the frontier), route-level variation in precision and jurisdiction, and a data-handling discipline that must be configured rather than assumed.
  6. Two things Anthropic does not publish decide the outcome: the numeric limits and how they vary. The monitor in recommendation 6 turns the calibrated estimates in this report into a live number.
10   Methodology, Limitations, Sources

How this was built

10.1 Method
  • Internal data. The principal's own telemetry, 2026-09-01 to 2026-09-15: a read-only scan of Claude Code transcript files (~/.claude/projects/**/*.jsonl, including subagent files), deduplicated by message, with token counts priced at Anthropic's published per-token rates (Opus $5/$25, Sonnet $2/$10, Haiku $1/$5, Fable $10/$50; cache reads at 0.1× input and writes at 1.25×).968 Active hours cluster the principal's own messages. Limit events were found by matching Claude Code's own limit notices and excluding 1,588 generic word matches and 12 self-referential research matches. The internal telemetry report (2026-09-08 to 2026-09-15) and the model bench are separate internal sources.6970
  • External data. Every vendor page was fetched on 2026-09-15 and each figure was read in the returned page text; JSON endpoints were used where available (OpenRouter's models and endpoints APIs). Pages that returned only JavaScript shells or HTTP 403 are marked and their figures are not used.
  • Derived figures. Every derived number in this report is computed by one script, with its formula recorded.
10.2 Limitations
  • Calibrated limits rest on two historical events on one account; Anthropic's limits are not published and may vary.
  • Every internal dollar figure is a Claude Code-only floor; the unmeasured share (0% / 15% / 30%) is an assumption, not a measurement.
  • The scenario model is linear in hours, holds token volume constant across model changes, estimates context-reset savings from growth curves rather than a replay, scales five-hour windows by the shape of observed days, and prices Fable work at Sonnet rates when computing mix savings (understating Fable savings).
  • The model-mix split between the usage pass (28% Fable) and the telemetry report (no Fable) is unreconciled; which Fable version ran is unknown.
  • The bench has 8–20 cases per suite, several constructed from rules, and does not cover xAI, most of Google, or any Chinese or open-weight model.
  • US-host prices move frequently; OpenRouter's listing is a snapshot of 2026-09-15 and reflects prices on OpenRouter, not necessarily each host's direct price.
  • No SOC 2 report, ISO certificate or data-processing agreement was inspected.
  • Statute descriptions are research, not legal advice; the Chinese text governs.
10.3 Claims dropped or left unverified
Claim in the input notesDisposition
Third-party estimates of Sonnet-hours and Opus-hours per week for Max tiers; Opus-to-Sonnet auto-fallback at 20% / 50%Dropped — not from Anthropic
"97,361 unique priced messages (25,512 since 9/1)"Corrected: 25,512 priced since 9/1; 97,361 is undated
Background share 32.8%Corrected to 33.2%
"Combined (1+3+4) −46%" described with the Opus → 15% leverCorrected: −46% uses Opus → 30%; Opus → 15% gives −53%
"Current (all-Opus)" scenario labelCorrected: current mix is 66.6% Opus, 28.0% Fable
Fable price UNVERIFIEDVerified $10 / $50 (Fable 5 cache reads $1.00; Fable 5.1 $0.25)
OpenAI flagship ~$10/$50, mid-tier ranges, gpt-5-mini (secondary)Replaced by verified OpenAI pages
ChatGPT Go $8 and Business seat prices; Codex headless flagDropped — not verified
Gemini CLI free tier "cut June 2026"Dropped — current docs list a free tier
xAI Heavy $300; Grok 4.3 $1.25 / $2.50Dropped — not in page text
Artificial Analysis index scores, Kimi K3 93.4% SWE-bench, "highest-ranked open-weights" claimsDropped — aggregator sources
Qwen Coder-Plus tier pricing, Qwen Flash via aggregatorDropped; Qwen3.8-Max verified instead
GLM Coding Plan Pro ~$72 and Max ~$160; MiniMax ¥9.9 planDropped — only "starting at 18 USD" verified
Doubao pricing; PRC phone verification requirementDropped — not readable / not verified
Model licences (DeepSeek MIT, Qwen Apache 2.0, Kimi "Modified MIT" and revenue threshold, GLM licence)Unverified — licence files not pulled
Z.ai excludes API content from training by defaultDropped — not in the privacy policy fetched
Federal Commerce and Navy DeepSeek bans; Texas, Virginia (order number conflicts between notes and search), Iowa, Kansas, South Dakota, Nebraska bansDropped — primary text not readable or not found; New York verified
H.R. 1121 current statusUnverified — introduced text verified only
BIS export-control rules (Jan 2026, May 31 2026 guidance); OMB M-24-10Dropped — secondary sources only
OpenRouter FedRAMP and HIPAA claims; Together, Groq, Baseten, Cerebras, SambaNova, Hyperbolic, Nvidia NIM, Cloudflare compliance claimsDropped or link-only — trust pages did not render; certificates not inspected
OpenRouter "data-retention filter"Contradicted by OpenRouter docs; zdr flag described instead
"Quantization rarely disclosed"Contradicted for OpenRouter endpoints, which label it
Fireworks spend-gate tiers; Together budget alertsDropped — not verified
AWS Bedrock and Vertex per-token prices; Cloudflare neurons pricing; NIM free tierDropped — not readable
Novita jurisdictionUnresolved
Local 8 GB VRAM throughput figuresDropped — secondary source
an internal notes file (named as an input)Not present in the input folder; not used
10.4 Sources

All external sources accessed 2026-09-15.

  1. Anthropic, "What is the Max plan?", Claude Help Center — https://support.claude.com/en/articles/11049741-what-is-the-max-plan Accessed 2026-09-15.
  2. Anthropic, "How do usage and length limits work?" — https://support.claude.com/en/articles/11647753-how-do-usage-and-length-limits-work Accessed 2026-09-15.
  3. Anthropic, "Use Claude Code with your Pro or Max plan" — https://support.claude.com/en/articles/11145838-use-claude-code-with-your-pro-or-max-plan Accessed 2026-09-15.
  4. Anthropic, "Manage usage credits for paid Claude plans" — https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans Accessed 2026-09-15.
  5. Anthropic, "Buy usage bundles" — https://support.claude.com/en/articles/14246112 Accessed 2026-09-15.
  6. Anthropic, "Claude Fable models on your plan" — https://support.claude.com/en/articles/15424964-claude-fable-models-on-your-plan Accessed 2026-09-15.
  7. Anthropic, "Models, usage, and limits in Claude Code" — https://support.claude.com/en/articles/14552983-models-usage-and-limits-in-claude-code Accessed 2026-09-15.
  8. Anthropic, "Use the Claude Agent SDK with your Claude plan" (update of June 15, 2026) — https://support.claude.com/en/articles/15036540 Accessed 2026-09-15.
  9. Anthropic, "Pricing," Claude Platform docs (model, cache, batch, data-residency and tokenizer notes) — https://platform.claude.com/docs/en/about-claude/pricing Accessed 2026-09-15.
  10. Anthropic, "Legal and compliance," Claude Code docs — https://code.claude.com/docs/en/legal-and-compliance Accessed 2026-09-15.
  11. Anthropic, Consumer Terms of Service — https://www.anthropic.com/legal/consumer-terms Accessed 2026-09-15.
  12. Anthropic, "Manage costs effectively," Claude Code docs — https://code.claude.com/docs/en/costs Accessed 2026-09-15.
  13. Anthropic, "Model configuration," Claude Code docs — https://code.claude.com/docs/en/model-config Accessed 2026-09-15.
  14. Anthropic, "Subagents," Claude Code docs — https://code.claude.com/docs/en/sub-agents Accessed 2026-09-15.
  15. Anthropic, "Status line," Claude Code docs — https://code.claude.com/docs/en/statusline Accessed 2026-09-15.
  16. OpenAI, API pricing — https://platform.openai.com/docs/pricing Accessed 2026-09-15.
  17. OpenAI, GPT-5 model page — https://platform.openai.com/docs/models/gpt-5 Accessed 2026-09-15.
  18. OpenAI, GPT-5 mini model page — https://platform.openai.com/docs/models/gpt-5-mini Accessed 2026-09-15.
  19. OpenAI, "What is ChatGPT Plus?" (includes the Pro $200 pause notice) — https://help.openai.com/en/articles/6950777-what-is-chatgpt-plus Accessed 2026-09-15.
  20. OpenAI, "Using Codex with your ChatGPT plan" — https://help.openai.com/en/articles/11369540-using-codex-with-your-chatgpt-plan Accessed 2026-09-15.
  21. Google, "Gemini Developer API pricing" — https://ai.google.dev/gemini-api/docs/pricing Accessed 2026-09-15.
  22. Google, Google AI subscriptions — https://gemini.google/subscriptions/ Accessed 2026-09-15.
  23. Google, "Everything new in our Google AI subscriptions, fresh from I/O 2026" — https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions/ Accessed 2026-09-15.
  24. Gemini CLI, "Quotas and pricing" — https://geminicli.com/docs/resources/quota-and-pricing/ (project repository: https://github.com/google-gemini/gemini-cli) Accessed 2026-09-15.
  25. xAI, "Grok Models & Pricing" — https://docs.x.ai/docs/models Accessed 2026-09-15.
  26. xAI, Pricing — https://x.ai/pricing Accessed 2026-09-15.
  27. xAI, "Introducing Grok Build" — https://x.ai/news/grok-build-cli Accessed 2026-09-15.
  28. xAI, Acceptable Use Policy — https://x.ai/legal/acceptable-use-policy Accessed 2026-09-15.
  29. DeepSeek, "Models & Pricing" — https://api-docs.deepseek.com/quick_start/pricing Accessed 2026-09-15.
  30. DeepSeek, Privacy Policy — https://cdn.deepseek.com/policies/en-US/deepseek-privacy-policy.html Accessed 2026-09-15.
  31. DeepSeek, Open Platform Terms of Service — https://cdn.deepseek.com/policies/en-US/deepseek-open-platform-terms-of-service.html Accessed 2026-09-15.
  32. Moonshot AI, Kimi API "Model Inference Pricing" — https://platform.kimi.ai/docs/pricing/chat Accessed 2026-09-15.
  33. Z.ai, Pricing — https://docs.z.ai/guides/overview/pricing Accessed 2026-09-15.
  34. Z.ai, GLM Coding Plan overview — https://docs.z.ai/devpack/overview Accessed 2026-09-15.
  35. Z.ai, Privacy Policy — https://docs.z.ai/legal-agreement/privacy-policy Accessed 2026-09-15.
  36. Alibaba Cloud, "Model Studio model pricing" — https://www.alibabacloud.com/help/en/model-studio/model-pricing Accessed 2026-09-15.
  37. MiniMax, "Pay as You Go Pricing" — https://platform.minimax.io/docs/guides/pricing-paygo Accessed 2026-09-15.
  38. BytePlus, ModelArk documentation (prices not readable) — https://docs.byteplus.com/en/docs/ModelArk/1099320 Accessed 2026-09-15.
  39. OpenRouter, public models API — https://openrouter.ai/api/v1/models Accessed 2026-09-15.
  40. OpenRouter, model endpoints API (example: DeepSeek V4 Pro) — https://openrouter.ai/api/v1/models/deepseek/deepseek-v4-pro/endpoints ; also /deepseek/deepseek-v4.1-flash, /moonshotai/kimi-k2.6, /moonshotai/kimi-k3, /z-ai/glm-5.3, /qwen/qwen3.8-max-0902, /openai/gpt-oss-120b, /minimax/minimax-m3 Accessed 2026-09-15.
  41. OpenRouter, "Provider selection" — https://openrouter.ai/docs/guides/routing/provider-selection Accessed 2026-09-15.
  42. OpenRouter, "Provider logging" — https://openrouter.ai/docs/guides/privacy/logging Accessed 2026-09-15.
  43. OpenRouter, "API credit and rate limits" — https://openrouter.ai/docs/api/reference/limits Accessed 2026-09-15.
  44. Fireworks AI, "Serverless pricing" — https://docs.fireworks.ai/serverless/pricing Accessed 2026-09-15.
  45. Fireworks AI, "Data security" — https://docs.fireworks.ai/guides/security_compliance/data_security Accessed 2026-09-15.
  46. DeepInfra, Pricing — https://deepinfra.com/pricing Accessed 2026-09-15.
  47. DeepInfra, Privacy Notice — https://deepinfra.com/privacy Accessed 2026-09-15.
  48. Amazon Web Services, "DeepSeek in Amazon Bedrock" — https://aws.amazon.com/bedrock/deepseek/ ; Bedrock pricing — https://aws.amazon.com/bedrock/pricing/ Accessed 2026-09-15.
  49. Microsoft, "Foundry Models sold directly by Azure" — https://learn.microsoft.com/en-us/azure/ai-foundry/foundry-models/concepts/models-sold-directly-by-azure Accessed 2026-09-15.
  50. Vercel, "AI Gateway pricing" — https://vercel.com/docs/ai-gateway/pricing ; model page — https://vercel.com/ai-gateway/models/deepseek-v4-flash Accessed 2026-09-15.
  51. Together AI, Trust Center (did not render) — https://trust.together.ai/ Accessed 2026-09-15.
  52. Groq, Trust Center (did not render) — https://trust.groq.com/ Accessed 2026-09-15.
  53. Baseten, Trust Center (did not render) — https://trust.baseten.co/ Accessed 2026-09-15.
  54. NIST, AI Risk Management Framework — https://www.nist.gov/itl/ai-risk-management-framework Accessed 2026-09-15.
  55. NIST, AI 600-1, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" — https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence Accessed 2026-09-15.
  56. NIST CAISI, "CAISI Evaluation of DeepSeek V4 Pro" (May 2026) — https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro Accessed 2026-09-15.
  57. NIST CAISI, "CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks" (Sept 2025) — https://www.nist.gov/news-events/news/2025/09/caisi-evaluation-deepseek-ai-models-finds-shortcomings-and-risks Accessed 2026-09-15.
  58. U.S. GPO GovInfo, H.R. 1121 (IH), 119th Congress — https://www.govinfo.gov/content/pkg/BILLS-119hr1121ih/html/BILLS-119hr1121ih.htm ; bill page (returned 403) — https://www.congress.gov/bill/119th-congress/house-bill/1121 Accessed 2026-09-15.
  59. Office of the Governor of New York, "Governor Hochul Issues Statewide Ban on DeepSeek Artificial Intelligence for Government Devices and Networks" (2025-02-10) — https://www.governor.ny.gov/news/governor-hochul-issues-statewide-ban-deepseek-artificial-intelligence-government-devices-and Accessed 2026-09-15.
  60. ISO, ISO/IEC 42001 (returned 403) — https://www.iso.org/standard/42001 Accessed 2026-09-15.
  61. AICPA & CIMA, "System and Organization Controls: SOC Suite of Services" — https://www.aicpa-cima.com/topic/audit-assurance/audit-and-assurance-greater-than-soc-2 Accessed 2026-09-15.
  62. GSA, FedRAMP — https://www.fedramp.gov/ ; Marketplace — https://marketplace.fedramp.gov/ Accessed 2026-09-15.
  63. EUR-Lex, Regulation (EU) 2024/1689 (Artificial Intelligence Act) — https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng Accessed 2026-09-15.
  64. National People's Congress of the PRC, Personal Information Protection Law (Chinese) — http://www.npc.gov.cn/npc/c2/c30834/202108/t20210820_313088.html Accessed 2026-09-15.
  65. National People's Congress of the PRC, Data Security Law (Chinese) — http://www.npc.gov.cn/npc/c2/c30834/202106/t20210610_311888.html Accessed 2026-09-15.
  66. National People's Congress of the PRC, National Intelligence Law, 2017 text (Chinese) — http://www.npc.gov.cn/zgrdw/npc/xinwen/2017-06/27/content_2024529.htm Accessed 2026-09-15.
  67. National People's Congress of the PRC, PIPL English page — http://www.npc.gov.cn/npc/c2597/c5854/bfflywwb/202311/t20231117_433007.html Accessed 2026-09-15.
  68. The principal's own telemetry, 2026-09-01 to 2026-09-15 — Claude Code transcript scan (usage, limit events, scenarios), method in §10.1
  69. The principal's own telemetry, 2026-09-08 to 2026-09-15 — internal report "Chat cost vs subscription," internal pricing reference
  70. Salud Capital internal model bench, 2026-09-15 — internal pricing reference
  71. internal source, internal pricing reference
About this report. A Real-World Test: Running Claude Max at $200 Without Running Out, and Where Other Vendors Fit. Salud Capital Research, September 15, 2026. ~9,900 words · 67 primary external sources, all accessed 2026-09-15. Vendor terms and prices change; every vendor statement is dated to its access date.