Published
Accuracy and Value: Why Luna Is the Everyday Choice for HealthTasks Agents
1 min read
HealthTasks keeps scoring frontier models inside Agents. For everyday use, GPT-5.6 Luna and GPT-6 Luna match the figures educators quote, at a fraction of the cost.
HealthTasks continues to monitor frontier models and how they perform inside the Agents harness. Every model gets the same prompt, the same program data, and the same tool loop. The question we keep asking is which model educators should use every day.
The recommendation
The best models for everyday use combine accuracy and value. In the latest HealthTasks Agents Benchmark, that is the OpenAI Luna series: GPT-5.6 Luna and GPT-6 Luna.
They perform on par with frontier models from Google and xAI, at a fraction of the cost. Reporting stays accurate and consistent. We did not see hallucinations or data errors. The figures match the evaluations the agent was asked to summarize.
Why that matters
A correct answer you can afford to run again is worth more than a richer answer you will ration. Luna lets faculty and program leaders use Agents through the term, for the questions that come up in meetings and CQI, without treating each run as a special occasion.
That is how educators maximize agent usage without sacrificing accuracy or quality.
What we will keep doing
We will keep scoring new frontier models in the same harness as they land. The recommendation moves when the evidence moves.
Scores, cost, method, and limitations are published with the benchmark:
HealthTasks Agents Benchmark, September 2026 →
For how we thought about model choice on earlier Agents runs, see Benchmarking AI models for educational analytics.
Get started
GPT-5.6 Luna and GPT-6 Luna are the models we recommend for everyday questions in HealthTasks Agents. If you want to see that workflow on your program's data, book a demo.