Engineering
GPT-6 Astra takes the lead on grounded video QA
A small change to how pipeline stages refer to objects put GPT-6 Astra clearly ahead of Gemini 3.7 Flash, at 28 times the cost.

In our previous post we compared Gemini 3.7 Flash with self-hosted Qwen3.8-27B on Perception Test's grounded video QA task. The main finding was that pipeline design mattered more than the model version. We have since added two more frontier models, GPT-6 Astra and Claude Opus 5.5, and the same finding showed up again.
On its first run, Astra scored 0.5675 HOTA, level with Gemini's 0.5694. The pipeline had been built around Gemini, and one stage depended on something Gemini does reliably and Astra does not: repeating an object's name word for word. When that stage switched to referring to objects by number, Gemini's score did not change, but Astra's rose to 0.6224. That makes Astra the first model we have tested that is clearly ahead of Gemini. It also shows that even the most capable models need a pipeline designed to play to their strengths and work around their weaknesses.
The short version: With object identity passed between stages by number, GPT-6 Astra scores 0.622 HOTA against 0.570 for Gemini 3.7 Flash (+0.053, p = 0.005). It picks the right object on 97.3% of questions and leads 8 of 10 question types. It also costs 28 times as much per run.
The setup
Nothing changed from the previous post except the models, a fix in stage 3 and a larger context window for Qwen, all described below. We used the same 300-question sample, the same four-stage pipeline (plan, keyframe, point, then carry with a SAM 3 tracker), the same prompts and the same official HOTA scorer. HOTA is the geometric mean of detection accuracy (did the system find the object?) and association accuracy (did the track stay on the same object over time?)
Limitations
- Every figure comes from the same 300-question sample used in the previous post, drawn from the Perception Test training split and stratified to match the validation split's categories. The comparison is consistent but narrow.
- Stage 3 samples each box again on every run, so part of each v3 difference is run-to-run variation.
- The v3 runs executed stage 4 on an RTX PRO 6000 Blackwell. The earlier Gemini and Qwen runs used an RTX 5090.
Why Astra's first score understated it
The pointing stage (stage 3) used to write a free-text label on every keyframe, and stage 4 grouped boxes into tracks by exact label. When a model reworded a label between keyframes, for example "the clear drinking glass" and then "clear drinking glass", the result was a second, false track for the same object. False tracks lower the association score even when every box is right.
Gemini almost never did this: it copied the stage-1 names on 99.6% of its boxes. For Gemini, the free-text handoff worked as designed, and referring to objects by number changed nothing. Astra reworded far more often. Rewording a name costs nothing when a person reads the output, but it breaks a stage that matches labels exactly.
| Model | Named → tracks, before | Named → tracks, v3 | Questions with extra tracks |
|---|---|---|---|
| GPT-6 Astra | 552 → 735 | 552 → 551 | 84 → 0 |
| Claude Opus 5.5 | 546 → 578 | 547 → 545 | 28 → 0 |
| Qwen3.8-27B | 544 → 577 | 540 → 535 | 26 → 0 |
| Gemini 3.7 Flash | 527 → 526 | 527 → 522 | 5 → 0 |
Before the fix, Astra produced 183 extra tracks spread over 84 of the 300 questions, which is the worst result across the board. To avoid that, stage 3 now names each box's object by its number in the stage-1 list instead of by a description, so each object keeps a single track. With that change, no model produces an extra track on any question. Where the track count is slightly below the named count, stage 3 left out an object it could not see.
Results
| Model | Before | v3 | Δ vs Gemini | 95% CI | p | Cost |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 0.5675 | 0.6224 | +0.053 | [+0.031, +0.077] | 0.005 | $200.42 |
| Claude Opus 5.5 | 0.5722 | 0.5944 | +0.025 | [+0.004, +0.046] | 0.21 | $82.41 |
| Gemini 3.7 Flash | 0.5694 | 0.5702 | +0.001 | [−0.005, +0.006] | 0.57 | $7.08 |
| Qwen3.8-27B | 0.4759 | 0.4807 | −0.089 | [−0.121, −0.058] | <0.001 | own GPUs |
HOTA before and after the v3 fix. Cost is the API bill for one full 300-question run. The 95% CI is a paired bootstrap of the mean per-question difference; p is from a paired Wilcoxon signed-rank test.
Astra gained the most from the fix (+0.055), because it had the most label drift to lose. Its interval is well clear of zero and the signed-rank test agrees. Opus's mean gain looks positive, but it comes from a small set of questions: against Gemini, Opus wins on 129 questions and loses on 115, and the signed-rank test does not separate them (p = 0.21). Gemini's score did not move.
Note that the previous post reported $9.29 for a 300-question Gemini 3.7 run, while the figure above is $7.08. The difference comes from another small optimization: stage 3 had been asking Gemini for segmentation masks, but stage 4 seeds the SAM 3 tracker from the box alone, so the masks were never used. Returning boxes only gave the same score with half the output tokens. This even further improves the already impressive performance per dollar score Gemini had.
Where Astra is strongest
The previous post found that most of the gap between models came from choosing the right object, not from drawing its box. Astra is best at that step: it picks the right object on 97.3% of questions, against 92.7% for Opus, 90.0% for Gemini and 85.0% for Qwen. When limited in how it can name objects, it also maintains track continuity.
| Question type | n | Gemini 3.7 Flash | Claude Opus 5.5 | GPT-6 Astra | Qwen3.8-27B |
|---|---|---|---|---|---|
| All questions | 300 | 0.570 | 0.594 | 0.622 | 0.481 |
| Cup games | 86 | 0.524 | 0.559 | 0.577 | 0.372 |
| Camera moving | 86 | 0.524 | 0.558 | 0.612 | 0.421 |
| 2-object answers | 29 | 0.617 | 0.648 | 0.679 | 0.553 |
| 4+ object answers | 23 | 0.508 | 0.501 | 0.543 | 0.399 |
| Memory questions | 17 | 0.436 | 0.455 | 0.407 | 0.453 |
| Abstraction questions | 13 | 0.662 | 0.673 | 0.638 | 0.535 |
Astra is the best model on 8 of the 10 slices we track. Its largest lead is on moving-camera clips, where it scores 0.612 against Gemini's 0.524. It also leads on cup games, the shell-game questions where the target is identifiable only from what happened earlier in the clip.
It has two weak spots. On memory questions Astra scores 0.407, the lowest of the four models, and on abstraction questions it is below both Opus and Gemini. Answers with four or more objects remain hard for every model. Astra is still best there, but its score drops to 0.543.
Was Qwen starved of input?
The v3 runs also tested a hypothesis about Qwen from the previous post: that its 32k-token context window was holding it back. We ran every stage again on a server with a 128k context and a larger thinking budget.
The larger window helped Qwen understand the questions: it picked the right object on 85.0% of questions, up from 81.7%. Very little of that reached the end-to-end score, which rose only from 0.4759 to 0.4807 and still trails Gemini by 0.089. Cup games remain its weakest area at 0.372, the lowest score in the table. Going to 128k also slowed Qwen down. Now we can tell that the context was not the main thing holding Qwen back.
The price of the lead
A 300-question run through the full pipeline costs $200.42 with Astra, $82.41 with Opus and $7.08 with Gemini. That is about 67 cents per question for Astra against a little over 2 cents for Gemini. Put another way, Astra's extra 0.053 HOTA costs about 64 cents per question. Per hour of video that is about $105 for Astra against $3.70 for Gemini, and our VLM tracker keeps these price and quality figures current as models and rates change.
Whether that is worth paying depends on what a wrong object costs. For offline audits, or for low-volume, high-stakes questions such as identifying which unit went through a faulty assembly step, Astra's object selection is worth paying for. For continuous monitoring across many cameras, Gemini still gives the most accuracy per dollar by a wide margin. The modular design from the previous post points to a hybrid: use the expensive model only for the stages where choosing the object is the bottleneck.
What no model fixes
Twenty-four questions score below 0.2 HOTA for all three API models, and 13 of those are cup games. The keyframe stage seeds the tracker only where the object is plainly visible, so an object that moves while hidden under a cup is not followed. This is a limit of the pipeline, not of any model, and switching models will not fix it.
Agnify answers questions like these from the cameras already on a production floor. To see the pipeline on your own footage, read about Agnify for manufacturing or book a demo.
FAQ
Which model is best at grounded video QA?
GPT-6 Astra, on this benchmark. With object identity passed between pipeline stages by number, it scores 0.622 HOTA against 0.570 for Gemini 3.7 Flash, a paired difference of +0.053 (p = 0.005). Claude Opus 5.5 scores 0.594 and Qwen3.8-27B 0.481. Astra also picks the right object most often, on 97.3% of questions.
Is GPT-6 Astra worth the cost over Gemini 3.7 Flash?
Only when a wrong object is expensive. A 300-question run costs $200.42 with Astra and $7.08 with Gemini, about 67 cents per question against a little over 2 cents. For offline audits and low-volume, high-stakes questions, Astra's better object selection is worth paying for. For continuous monitoring across many cameras, Gemini gives far more accuracy per dollar.
Why did Astra score lower on its first run?
Its first run scored 0.5675 HOTA, level with Gemini. The pointing stage wrote a free-text label on every keyframe, and the tracking stage grouped boxes into tracks by exact label. Astra often reworded a label between keyframes, which split one object into two tracks: 183 extra tracks across 84 of the 300 questions. Once the stage referred to objects by number instead, the extra tracks disappeared and Astra's score rose to 0.6224.
How does Claude Opus 5.5 compare?
Opus scores 0.594 HOTA, 0.025 above Gemini on average, but the gain is not statistically significant (p = 0.21): against Gemini it wins on 129 questions and loses on 115. A 300-question run costs $82.41 with Opus against $7.08 with Gemini.
Does a 128k context window help Qwen3.8-27B?
Barely. With a 128k context and a larger thinking budget, Qwen picked the right object on 85.0% of questions, up from 81.7%, but its HOTA rose only from 0.4759 to 0.4807 and still trails Gemini by 0.089. The larger window also made it slower, so context was not the main thing holding Qwen back.



