HealthTasks Research · v0.1 · 26 Sep 2026

HealthTasks Agents Benchmark

This benchmark asks a practical question: does a model have to lead a general frontier benchmark to be the right model inside the HealthTasks Agents harness? Every model here gets the same prompt, the same dataset, and the same tool loop. Accuracy is checked against the source evaluations. Cost is what it takes to run the job again.

Accuracy and cost are the two results that decide how much intelligence an institution actually gets. A correct answer you can afford to repeat is worth more than a slightly richer answer you will ration. The Luna models stay with the other frontier models, and with more expensive models in their own family, at a fraction of the cost. That is the argument of this benchmark: educators get more from their agents when the accurate run is the inexpensive one.

1
task
119
evaluations
6
models
7
runs

Leaderboard

sorted by task score · gateway cost summed

#ModelTaskDepthCostOut tokStepsTime
1
gpt-5.6-luna [max]
OpenAI
100
0
$0.0294.2k753.2s
2
gemini-3.8-flash
Google
97
30
$0.2533.7k961.6s
3
grok-4.6
xAI
85
0
$0.2515.8k675.0s
4
gpt-6-luna [high]
OpenAI
82
0
$0.009–0.0123.1–7.5k634–55s
5
gpt-6-sol
OpenAI
67
0
$0.2314.6k465.0s
6
grok-4.7
xAI
10
0
$0.1501.7k426.4s

Task is 100 points on the prompt. Depth is 30 extra points and does not change the rank. Input tokens: Luna 5.6 [max] 207,702 · Gemini 319,053 · Grok 4.6 193,188 · Luna 6 [high] 151,565 · Sol 100,172 · Grok 4.7 115,319.

Task

The agent receives a fixed institution in the HealthTasks development environment and the prompt below. It may query submitted evaluations, write a short answer, and attach a canvas.

evaluation trends over time

A full task score reports 119 scored submissions, 94.11%, and the monthly series, with inactive months left as gaps. Classroom and template splits add depth. They do not raise the task score.

Method continues the July educational-analytics bakeoff and the June clinical-analytics bakeoff. Those runs compared two models. This v0.1 bakeoff holds the harness fixed and swaps the model.

Ground truth

Submissions119
Overall average94.11%
April 202543 · 96.28%
June 202514 · 84.02%
September 20269 · 97.04%
ICU78 · 95.72%
Med Surg32 · 90.42%
Unassigned8 · 92.50%
Classroom “1”1 · 100%
Daily Eval43 · 94.26%
Preceptor34 · 92.94%
Location32 · 98.13%

Scoring

ScoreCriterionPtsRule
TaskMonthly series50Checked months match on count and mean = 50. Correct quarter totals only = 20. Any stated month with the wrong volume = 0.
TaskOverall mean3094.11% = 30. Correct one-decimal 94.1% = 27. Off by 0.1 = 15. Omitted = 12. Off by more than 0.1 = 0.
TaskSubmission count20119 scored submissions = 20. Stating 119 while treating some rows as unscored = 10.
DepthClassroom split15ICU, Med Surg, Unassigned, and classroom “1” match the source. Not required by the prompt.
DepthTemplate split15Daily Eval, Preceptor, and Location match the source. Not required by the prompt.

Notes

  • gpt-5.6-luna [max]

    task 100 · depth 0 · $0.029

    Exact 94.11% and the exact monthly series, with inactive months left as gaps. No classroom or template split.

  • gemini-3.8-flash

    task 97 · depth 30 · $0.253

    Monthly series matched. Headline rounded to 94.1%. Only run with a correct classroom split and a correct template split.

  • grok-4.6

    task 85 · depth 0 · $0.251

    Monthly chart and counts matched. Headline reported as 94.2%.

  • gpt-6-luna [high]

    task 82 · depth 0 · $0.009–0.012

    Two [high] runs both scored 82. Every month matched the source. Both omitted 94.11%. The first put volume in the axis labels. The second added a bar chart of submissions next to the score line. No classroom or template split. The later run was 34.3s and $0.009.

  • gpt-6-sol

    task 67 · depth 0 · $0.231

    Quarterly means and counts matched, and the headline is 94.1%. Month grain is omitted.

  • grok-4.7

    task 10 · depth 0 · $0.150

    Stated 119 evaluations, then averaged 115 of them to 94.4%. All 119 source rows are scored, at 94.11%. Several month volumes are wrong, including April 2025 at 41 instead of 43. Classroom counts do not match.

Limitations

  • v0.1 is one task. A model that wins here can lose on a different Agents prompt.
  • GPT-6 Luna [high] has two finished runs, both task 82. The other models have one finished run. That is still a small sample, not a confidence interval.
  • GPT-5.6 Luna [max] is the finished 5.6 run. A GPT-6 Luna [max] run was cancelled after it looped.
  • GPT-6 Luna [high] and Grok 4.7 were checked on all 15 active months. Earlier monthly credit uses the series the run displayed plus the April 2025, June 2025, and September 2026 anchors.
  • Cost and latency include the tool loop, not a single completion.

Conclusions

The HealthTasks Agents harness does not need the bleeding edge of a public leaderboard to land in the same place. GPT-5.6 Luna [max] leads this task at 100 for $0.029. Gemini 3.8 Flash scores 97 and is the only run with extra depth, at $0.253. GPT-6 Luna [high] matched the monthly series on two separate runs, at $0.009 to $0.012. Grok 4.6 and GPT-6 Sol cost about a quarter and do not beat Luna on the answer an educator would repeat. More usage is more intelligence. The models that keep the figures right at a fraction of the cost are the ones that let educators run the agent more often, and get more from it.

ModelRecommend forAgainst the other options
gpt-5.6-luna [max]Accuracy and valueTask 100 at $0.029. It states the figure exactly and stays with Gemini’s task score at about one-eighth the cost. Use it when the number will be quoted and the job should be cheap enough to run again.
gpt-6-luna [high]ValueTwo runs, both task 82, at $0.009–$0.012. The series matched both times. That is consistent accuracy at roughly twenty times below Gemini, Grok 4.6, and Sol. The headline average was left out, so use [max] when that number has to be written down.
gemini-3.8-flashExtra depthTask 97, depth 30, at $0.253. The extra cuts are real. They are not required to arrive at a correct answer, and they cost about eight times Luna [max]. Use Gemini when those cuts are the product, not as the default.
grok-4.6NeitherA usable series and a headline a tenth high. Task 85 at $0.251, in Gemini’s cost band, without Gemini’s depth and without Luna’s price. Luna is the better accuracy and value choice.
gpt-6-solNeitherSame family as Luna, coarser answer, much higher cost. Task 67 at $0.231. Luna [max] and Luna [high] both get further on the series for a fraction of that spend.
grok-4.7NeitherNeither accuracy nor value. It reported 94.4% on 115 rows; the source is 94.11% on 119. April 2025 was 41 instead of 43, and the classroom counts were wrong. Task 10 at $0.150.