HealthTasks Research · v0.1 · 26 Sep 2026
HealthTasks Agents Benchmark
This benchmark asks a practical question: does a model have to lead a general frontier benchmark to be the right model inside the HealthTasks Agents harness? Every model here gets the same prompt, the same dataset, and the same tool loop. Accuracy is checked against the source evaluations. Cost is what it takes to run the job again.
Accuracy and cost are the two results that decide how much intelligence an institution actually gets. A correct answer you can afford to repeat is worth more than a slightly richer answer you will ration. The Luna models stay with the other frontier models, and with more expensive models in their own family, at a fraction of the cost. That is the argument of this benchmark: educators get more from their agents when the accurate run is the inexpensive one.
- 1
- task
- 119
- evaluations
- 6
- models
- 7
- runs
Leaderboard
sorted by task score · gateway cost summed
| # | Model | Task | Depth | Cost | Out tok | Steps | Time |
|---|---|---|---|---|---|---|---|
| 1 | gpt-5.6-luna [max] OpenAI | 100 | 0 | $0.029 | 4.2k | 7 | 53.2s |
| 2 | gemini-3.8-flash Google | 97 | 30 | $0.253 | 3.7k | 9 | 61.6s |
| 3 | grok-4.6 xAI | 85 | 0 | $0.251 | 5.8k | 6 | 75.0s |
| 4 | gpt-6-luna [high] OpenAI | 82 | 0 | $0.009–0.012 | 3.1–7.5k | 6 | 34–55s |
| 5 | gpt-6-sol OpenAI | 67 | 0 | $0.231 | 4.6k | 4 | 65.0s |
| 6 | grok-4.7 xAI | 10 | 0 | $0.150 | 1.7k | 4 | 26.4s |
Task is 100 points on the prompt. Depth is 30 extra points and does not change the rank. Input tokens: Luna 5.6 [max] 207,702 · Gemini 319,053 · Grok 4.6 193,188 · Luna 6 [high] 151,565 · Sol 100,172 · Grok 4.7 115,319.
Task
The agent receives a fixed institution in the HealthTasks development environment and the prompt below. It may query submitted evaluations, write a short answer, and attach a canvas.
evaluation trends over time
A full task score reports 119 scored submissions, 94.11%, and the monthly series, with inactive months left as gaps. Classroom and template splits add depth. They do not raise the task score.
Method continues the July educational-analytics bakeoff and the June clinical-analytics bakeoff. Those runs compared two models. This v0.1 bakeoff holds the harness fixed and swaps the model.
Ground truth
| Submissions | 119 |
|---|---|
| Overall average | 94.11% |
| April 2025 | 43 · 96.28% |
| June 2025 | 14 · 84.02% |
| September 2026 | 9 · 97.04% |
| ICU | 78 · 95.72% |
| Med Surg | 32 · 90.42% |
| Unassigned | 8 · 92.50% |
| Classroom “1” | 1 · 100% |
| Daily Eval | 43 · 94.26% |
| Preceptor | 34 · 92.94% |
| Location | 32 · 98.13% |
Scoring
| Score | Criterion | Pts | Rule |
|---|---|---|---|
| Task | Monthly series | 50 | Checked months match on count and mean = 50. Correct quarter totals only = 20. Any stated month with the wrong volume = 0. |
| Task | Overall mean | 30 | 94.11% = 30. Correct one-decimal 94.1% = 27. Off by 0.1 = 15. Omitted = 12. Off by more than 0.1 = 0. |
| Task | Submission count | 20 | 119 scored submissions = 20. Stating 119 while treating some rows as unscored = 10. |
| Depth | Classroom split | 15 | ICU, Med Surg, Unassigned, and classroom “1” match the source. Not required by the prompt. |
| Depth | Template split | 15 | Daily Eval, Preceptor, and Location match the source. Not required by the prompt. |
Notes
gpt-5.6-luna [max]
task 100 · depth 0 · $0.029
Exact 94.11% and the exact monthly series, with inactive months left as gaps. No classroom or template split.
gemini-3.8-flash
task 97 · depth 30 · $0.253
Monthly series matched. Headline rounded to 94.1%. Only run with a correct classroom split and a correct template split.
grok-4.6
task 85 · depth 0 · $0.251
Monthly chart and counts matched. Headline reported as 94.2%.
gpt-6-luna [high]
task 82 · depth 0 · $0.009–0.012
Two [high] runs both scored 82. Every month matched the source. Both omitted 94.11%. The first put volume in the axis labels. The second added a bar chart of submissions next to the score line. No classroom or template split. The later run was 34.3s and $0.009.
gpt-6-sol
task 67 · depth 0 · $0.231
Quarterly means and counts matched, and the headline is 94.1%. Month grain is omitted.
grok-4.7
task 10 · depth 0 · $0.150
Stated 119 evaluations, then averaged 115 of them to 94.4%. All 119 source rows are scored, at 94.11%. Several month volumes are wrong, including April 2025 at 41 instead of 43. Classroom counts do not match.
Limitations
- v0.1 is one task. A model that wins here can lose on a different Agents prompt.
- GPT-6 Luna [high] has two finished runs, both task 82. The other models have one finished run. That is still a small sample, not a confidence interval.
- GPT-5.6 Luna [max] is the finished 5.6 run. A GPT-6 Luna [max] run was cancelled after it looped.
- GPT-6 Luna [high] and Grok 4.7 were checked on all 15 active months. Earlier monthly credit uses the series the run displayed plus the April 2025, June 2025, and September 2026 anchors.
- Cost and latency include the tool loop, not a single completion.
Conclusions
The HealthTasks Agents harness does not need the bleeding edge of a public leaderboard to land in the same place. GPT-5.6 Luna [max] leads this task at 100 for $0.029. Gemini 3.8 Flash scores 97 and is the only run with extra depth, at $0.253. GPT-6 Luna [high] matched the monthly series on two separate runs, at $0.009 to $0.012. Grok 4.6 and GPT-6 Sol cost about a quarter and do not beat Luna on the answer an educator would repeat. More usage is more intelligence. The models that keep the figures right at a fraction of the cost are the ones that let educators run the agent more often, and get more from it.
| Model | Recommend for | Against the other options |
|---|---|---|
| gpt-5.6-luna [max] | Accuracy and value | Task 100 at $0.029. It states the figure exactly and stays with Gemini’s task score at about one-eighth the cost. Use it when the number will be quoted and the job should be cheap enough to run again. |
| gpt-6-luna [high] | Value | Two runs, both task 82, at $0.009–$0.012. The series matched both times. That is consistent accuracy at roughly twenty times below Gemini, Grok 4.6, and Sol. The headline average was left out, so use [max] when that number has to be written down. |
| gemini-3.8-flash | Extra depth | Task 97, depth 30, at $0.253. The extra cuts are real. They are not required to arrive at a correct answer, and they cost about eight times Luna [max]. Use Gemini when those cuts are the product, not as the default. |
| grok-4.6 | Neither | A usable series and a headline a tenth high. Task 85 at $0.251, in Gemini’s cost band, without Gemini’s depth and without Luna’s price. Luna is the better accuracy and value choice. |
| gpt-6-sol | Neither | Same family as Luna, coarser answer, much higher cost. Task 67 at $0.231. Luna [max] and Luna [high] both get further on the series for a fraction of that spend. |
| grok-4.7 | Neither | Neither accuracy nor value. It reported 94.4% on 115 rows; the source is 94.11% on 119. April 2025 was 41 instead of 43, and the classroom counts were wrong. Task 10 at $0.150. |