Professional Agent Reliability Brief

Current observation window

Documented access is not the same as work that finishes.

· A buyer brief on completion trust, quota/accounting friction, and workflow interruption across Claude Code, Codex, and Cursor.

Three decision conclusions · not ranked

What the current receipts support—and what they do not.

Each conclusion uses the same evidence contract projected into provider and compare verdicts. Follow the numbered paths to audit it.

Claude Code

Strong model and workflow affinity remains visible, but current receipts bound completion trust and usage-control confidence.

Bounded caution

Buyer implication: Keep Claude Code on the shortlist for its workflow fit; do not assume remaining quota guarantees an uninterrupted session or completed agent work.

Confidence
Moderate · corroborated reports
Evidence
first-party issue report · practitioner report
Corroboration
Two independent Jul 17 reports describe a usage-credit interruption; a separate Jul 12 issue pairs model praise with completion and usage-burn failure.
Limits / contrary evidence
Reports are self-selected and do not establish prevalence. The completion report explicitly describes the model as better and cleaner.
Window · revalidate

Codex

Capability and a concrete switching account are cautiously positive; quota accounting, capacity stops, and Windows app incidents limit completion trust.

Bounded caution

Buyer implication: Trial Codex for the work it appears to fit, but verify quota telemetry and Desktop stability on your model, plan, and operating system before switching.

Confidence
Moderate · corroborated reports
Evidence
first-party issue report · community incident report · practitioner report
Corroboration
Separate reports cover fast quota drain, delayed or disputed accounting, capacity stops, and repeated Windows hangs; one practitioner reports switching direction for a specific use case.
Limits / contrary evidence
No receipt measures prevalence. Windows incidents do not establish CLI or macOS reliability, and one switching account is not market momentum.

Cursor

A supported access-interruption warning exists; current evidence is insufficient for cross-billing or broad quality-decline claims.

Insufficient current evidence

Buyer implication: Treat invoice-state interruption as a workflow risk. Do not change providers on the cross-billing allegation without account-level corroboration.

Confidence
Limited · isolated or attribution-limited
Evidence
reported incident coverage · community incident report
Corroboration
The access interruption has secondary coverage; the cross-billing claim has only one unresolved direct forum report in this window.
Limits / contrary evidence
The billing attribution is unresolved, no broad quality sample is present, and neither receipt supports prevalence or causation.
Window · revalidate

Direct receipt ledger

The claims, dates, and attribution limits.

Grouped by provider, not counted against one another. Receipt volume is not prevalence.

Claude Code

1

· first-party issue report

Claude Code issue #76987 ↗

One reporter described strong model output but repeated usage burn on re-reads and process that did not complete the requested work.

Attribution limit
Single self-reported workflow; the issue does not establish prevalence or a platform-wide completion rate.
Accessed
2

· first-party issue report

Claude Code issue #78613 ↗

A Max subscriber reported a selected model requiring usage credits two days before the transition date they understood had been communicated.

Attribution limit
Reporter interpretation of rollout timing; not an Anthropic incident confirmation and not evidence about every model or account.
Accessed
3

· practitioner report

Kyle Little on X ↗

A practitioner reported an in-session usage-credits requirement while quota remained.

Attribution limit
Single public post; account state and root cause are not independently verified.
Accessed

Codex

1

· first-party issue report

Codex issue #32827 ↗

One Desktop user documented model-attribution mismatch and delayed seven-day quota updates after heavy work.

Attribution limit
Single account with local telemetry; internal routing and billed attribution were not confirmed by OpenAI.
Accessed
2

· first-party issue report

Codex issue #32606 ↗

A Windows user reported selected models exhausting a five-hour usage window and purchased credits within minutes during repository work.

Attribution limit
One configuration and reasoning level; it does not quantify normal consumption across plans or models.
Accessed
3

· community incident report

OpenAI community capacity thread ↗

A Desktop user reported active tasks repeatedly stopping with “Selected model is at capacity.”

Attribution limit
Community report, not an OpenAI status incident; affected population and duration are unknown.
Accessed
4

· first-party issue report

Codex issue #33873 ↗

A Windows 10 user documented repeated Codex Desktop hangs after an update, including Windows application-hang events.

Attribution limit
One Windows installation; not evidence about CLI reliability or other operating systems.
Accessed
5

· practitioner report

Akash on X ↗

One practitioner said Codex worked better than Claude Code for their use case after using it to verify Claude-produced work.

Attribution limit
One use case and self-reported switch direction; it does not support market-wide switching momentum.
Accessed

Cursor

1

· reported incident coverage

Digg invoice-pause report ↗

Coverage described a recurring unpaid-invoice pause bug interrupting Cursor access.

Attribution limit
Secondary coverage of user reports; affected population and recurrence rate are not established.
Accessed
2

· community incident report

Cursor forum cross-billing thread ↗

One user reported Anthropic usage appearing while Claude was disabled in Cursor.

Attribution limit
Single unresolved forum report; routing, account linkage, and whether Cursor caused the usage are not established.
Accessed

Method / linking

A bounded weekly evidence contract.

We preserve each source’s unit of claim and refuse prevalence, causation, or market-wide switching inference. Revalidate on .