Evaluation report

The results.
And their limits.

Inspect what we tested, which model we used, and what the results do—and do not—establish.

GPT-5.6 Luna validation pending. The historical task results below use GPT-5.4, not the current production route. DeepSWE is a separate archived coding run.

Historical results ·

Task relevance

14/14 investigated cases passing all checks

Synthetic development testgpt-5.4-2026-03-05

Small, agent-authored fixtures—not real-world task accuracy.

Engineering enrichment

24/24 correct readiness decisions

Synthetic development testgpt-5.4-2026-03-05 · low reasoning

Readiness decisions—not independently reviewed plan quality.

Datacurve DeepSWE

64/113 full-cohort tasks passed (56.6%)

Separate Compress Cloud runArchived Compress Cloud coding run

Separate external coding benchmark—not a customer PR success rate or leaderboard comparison.

Separate measurements of task relevance, plan preparation and coding. Development fixtures are not independently labeled real-world accuracy. These scores must not be combined into an end-to-end success rate.

Task relevance

investigated cases passing all checks
14/14
Real tasks retained
14/14
Incorrect queue admissions
0/14
Source-only baseline passing all checks
8/14
  • 14 agent-authored cases; live reviewer with frozen history, code and web retrieval. Planning is an invocation spy, not an executed plan.
  • A real task may remain visible without belonging in our engineering queue. Human acceptance and execution approval remain separate.

Engineering enrichment

correct readiness decisions
24/24
Completed cases
24/24
False-ready decisions
0/12
Automatic wording checks
23/24
Plain-model readiness baseline
24/24
Plain-model wording baseline
20/24
Raw reviewer contrast checks
8/10
  • 24 agent-authored ambiguous tasks (16 original + 8 transfer). Readiness is not plan correctness. The separate raw reviewer scored 8/10, incorrectly rejecting two valid briefs.
  • Same source evidence and per-call output cap; unequal total budget: up to 4 Actionairy calls versus 2 baseline calls. No claim of competitor superiority.
  • Independently reviewed usable-plan accuracy and human correction time are not yet measured.
  • Failed calls remain in the full denominator. Automatic wording checks are lexical, not semantic grading. No independent human usable-plan score is available.

Datacurve DeepSWE

full-cohort tasks passed
64/113
  • 2 tasks remained ungradeable after bounded regrading. They receive no pass credit and stay in the full-cohort denominator.
  • Sealed candidate patches and pinned verifier contexts. This is a separate coding evaluation, not an end-to-end Actionairy meeting-to-code test or a current-model leaderboard claim.

What is not measured yet

Meeting extraction: independently labeled capture precision and recall are not yet available. Transcript-upload acceptance is a user-journey check, not an accuracy estimate.

Independent usable-plan accuracy, human correction time, and end-to-end voice-to-PR success are not yet measured. These results do not establish competitor superiority.

Provenance and attribution

The downloadable snapshot includes source commit identifiers and report hashes. It contains aggregate results, not private customer conversations.

The DeepSWE mark identifies Datacurve’s benchmark. Our reported result is a Light Reach evaluation, not a Datacurve endorsement or an official leaderboard entry.