Task relevance
14/14 investigated cases passing all checks
Small, agent-authored fixtures—not real-world task accuracy.
Inspect what we tested, which model we used, and what the results do—and do not—establish.
GPT-5.6 Luna validation pending. The historical task results below use GPT-5.4, not the current production route. DeepSWE is a separate archived coding run.
14/14 investigated cases passing all checks
Small, agent-authored fixtures—not real-world task accuracy.
24/24 correct readiness decisions
Readiness decisions—not independently reviewed plan quality.
64/113 full-cohort tasks passed (56.6%)
Separate external coding benchmark—not a customer PR success rate or leaderboard comparison.
Separate measurements of task relevance, plan preparation and coding. Development fixtures are not independently labeled real-world accuracy. These scores must not be combined into an end-to-end success rate.
Meeting extraction: independently labeled capture precision and recall are not yet available. Transcript-upload acceptance is a user-journey check, not an accuracy estimate.
Independent usable-plan accuracy, human correction time, and end-to-end voice-to-PR success are not yet measured. These results do not establish competitor superiority.
The downloadable snapshot includes source commit identifiers and report hashes. It contains aggregate results, not private customer conversations.
The DeepSWE mark identifies Datacurve’s benchmark. Our reported result is a Light Reach evaluation, not a Datacurve endorsement or an official leaderboard entry.