We use public video retrieval datasets to benchmark search quality. The setup looks simple: take a text query, retrieve video moments, and compare the result against the dataset's ground-truth clip.

The benchmark score depends on one assumption.

The dataset ground truth has to be right.

That assumption does not always hold. During manual review, we found cases across MSRVTT, MSVD, VATEX, DiDeMo, and QVHighlights where the dataset annotation looked wrong, too narrow, or too hard to defend as the only valid answer.

Some retrieved clips matched the query visually, but they sat outside the accepted ground-truth window. In a benchmark report, those results look like retrieval misses. When you watch the clips, some of them look like dataset misses.

For example:

  • A query like a car is in a wreck can match several crash moments, not just one timestamp.
  • A query like a video game is played is broad enough to describe many gameplay clips.
  • A query like A woman is doing a big hole in a pumpkin. can match the carving process before the annotated window.
  • A query like Most beautiful resort i have ever seen depends on subjective context, not a single visual fact.

This matters because a retrieval score can mix two things: how well the system finds the right moment, and how well the dataset represents the possible right moments.

The second part is easy to miss when we only look at aggregate metrics.

We saw five annotation patterns:

  • Narrow windows that capture only part of an event.
  • Shifted timestamps where matching evidence appears nearby.
  • Missing alternate positives when more than one clip is valid.
  • Broad captions that allow several plausible clips.
  • Subjective captions that depend on context or judgment.

Methodology

We reviewed the query, the ground-truth clip, and the top retrieved shots side by side.

A case was included only when the retrieved clips had visible evidence matching the query. The useful question in those cases was whether the dataset annotation was complete and correct.

We left out examples that looked better explained by search quality, indexing, clip localization, or retrieved-shot descriptions with limited visual detail.

This was a qualitative review. The goal was to understand where one accepted timestamp can hide reasonable visual alternatives.

Annotation patterns

The examples below use the same vocabulary throughout the note.

Dataset issue What it looks like
Narrow GT window The accepted clip captures only a small part of the event.
Shifted GT timestamp Relevant evidence appears before or after the accepted window.
Missing alternate positives Other clips satisfy the query, but the benchmark accepts one answer.
Broad query with one target The query supports multiple valid clips, while scoring expects one.
Subjective caption The query depends on judgment, such as beauty or context.

Clip evidence

Each example shows the accepted GT clip first, followed by the top 5 retrieved shots.

The clips make the dataset issue easier to inspect.

MSRVTT: a car is in a wreck

  • GT: on video9800
  • Many returned segments clearly show car crashes or wrecks, but they fall outside or partially outside the annotated GT window. The GT appears narrower than the set of semantically matching segments.
Rank Window Clip Why it matters
1 0.0-5.0s A red rally car speeds along a winding paved road, loses control, and slides off the road into a grassy embankment.
2 20.0-25.0s A yellow-and-blue rally car skids off a tree-lined road and flips multiple times.
3 30.0-40.0s A rally car loses control, slides off the road, and kicks up tire smoke.
4 45.0-80.0s A yellow rally car drifts through a turn, skids off course, and crashes into roadside protection.
5 115.0-120.0s A rally car overturns on a bend and lands upside down at the edge of the road.

MSRVTT: a video game is played

  • GT: on video8027
  • Many returned scenes explicitly show video games being played across multiple timestamps. The query is generic, so one accepted GT window captures only part of the visual concept.
Rank Window Clip Why it matters
1 525.225-530.23s A game screen displays a round-end message and visible gameplay interface.
2 540.24-545.245s A game character fires a projectile across a dark arena with scores and time on screen.
3 550.25-555.255s A game character attacks enemies in a grid-patterned arena.
4 560.26-565.265s Arcade-style gameplay continues with a player character moving and firing.
5 595.295-600.3s A player character navigates a game screen with enemies and interface elements.

MSVD: a delicious japanese dish

  • GT: on -wa0umYJVGg_100_115
  • The candidate moments show several different Japanese-food preparation scenes, but the benchmark accepts only a narrow annotated answer for a broad and subjective caption.
Rank Window Clip Why it matters
1 13.012-17.0s Onigiri is arranged with bento sides.
2 5.005-13.012s A battered piece of food, likely tempura, is turned with chopsticks while frying in bubbling oil.
3 0.0-5.004s Hands shape white rice around cooked salmon to make a filled onigiri.
4 30.0-40.0s A cutlet is rolled and pressed into a sesame or crumb coating.
5 45.0-50.0s The cutlet is flipped and pressed through breadcrumbs.

VATEX: A woman is doing a big hole in a pumpkin.

  • GT: on efvOYBo03XM_000351_000361
  • Multiple returned clips are clear semantic matches for making a big hole in a pumpkin, but they occur outside the annotated GT window. The action appears to span more than the accepted timestamp.
Rank Window Clip Why it matters
1 60.06-65.065s A young woman stands beside a large pumpkin and begins working on it.
2 65.065-70.07s A woman is focused on carving a large orange pumpkin at a table.
3 75.075-100.1s The same pumpkin-carving activity continues in a longer window.
4 100.1-115.115s The person continues actively carving the pumpkin on the table.
5 115.115-125.125s A carving tool is used to saw into the large pumpkin.

DiDeMo: head passes in front of camera

  • GT: on 53301297@N00_5826898997_0a951bea4f
  • Returned clips show heads or faces moving into the immediate foreground at other timestamps. The query is broad, and the accepted GT captures one instance of a repeated visual event.
Rank Window Clip Why it matters
1 40.04-45.045s The camera moves suddenly from people in the street toward the foreground.
2 54.054-57.057s A person in a Santa hat moves their face closer to the camera.
3 0.0-5.005s A man leans close until his face fills the frame.
4 48.048-51.051s A man approaches from a doorway and leans into the camera.
5 18.018-21.021s A close face leaves the frame as the camera tilts upward.

DiDeMo: castle comes into view

  • GT: on 51167579@N06_6829893951_f10a25a3c8
  • Several returned windows show the same semantic target, a tower or castle-like structure entering view, but the benchmark accepts a narrow GT span.
Rank Window Clip Why it matters
1 69.0-77.2s A camera pan reveals a tall, dark church or tower structure.
2 5.005-10.01s The camera moves along a path toward distant stone architecture.
3 12.012-20.02s A stone tower remains visible while the viewpoint moves along the path.
4 39.039-42.042s A slow zoom brings the stone church tower closer in a churchyard view.
5 20.02-25.025s The viewpoint reveals more of the stone tower behind trees.

DiDeMo: light starts blue

  • GT: on 26292851@N04_4223864218_8a531d1c08
  • The accepted GT is subtle, while top-ranked clips contain clearer blue-light events outside the GT. The query is visually under-specified for a single narrow target.
Rank Window Clip Why it matters
1 30.03-35.035s Blue lights appear and move across the airport tarmac view.
2 40.04-45.045s A line of blue runway lights appears from an aircraft-window view.
3 70.07-75.075s Bright ground lights pass through the frame during aircraft movement.
4 0.0-5.005s A silhouetted rower moves through glowing blue cave water.
5 5.005-15.015s A performer stands on an outdoor stage under blue light and smoke.

QVHighlights: A girl speaking from her car

  • GT: on Zhx9Ki9bUkE_360.0_510.0
  • The GT is only about two seconds, while the video contains many car-speaking moments that satisfy the query.
Rank Window Clip Why it matters
1 45.045-60.06s A young woman sits in the driver's seat of a car and speaks to camera.
2 105.105-120.12s A young woman speaks inside a car at night.
3 150.15-165.165s A blonde woman with glasses speaks from the driver's seat.
4 450.15-465.165s A conversational car scene shows a young woman and an older man.
5 120.12-135.135s A young woman sits in a car backseat and speaks directly to the camera.

QVHighlights: Most beautiful resort i have ever seen

  • GT: on lyGaTk4MLVM_60.0_210.0
  • The caption is subjective, and many resort scenes can satisfy it outside the small annotated GT windows.
Rank Window Clip Why it matters
1 255.0-270.0s A panoramic beachfront resort view transitions into a hotel-room view.
2 285.0-315.0s A guided room tour shows a luxury room and its outdoor surroundings.
3 750.0-765.0s The camera moves through a lush tropical resort toward a spa entrance.
4 795.0-810.0s A relaxing tropical resort or spa scene appears.
5 840.0-855.0s The segment shows a resort or spa vacation setting.

Findings

The strongest examples followed a few repeatable annotation patterns.

Many annotations were wrong because they were too narrow for repeated or extended actions. Car wrecks, gameplay, pumpkin carving, and car-speaking moments can occur across many windows in the same or similar videos.

A single GT span can be brittle for those captions.

Generic captions create multiple plausible answers. Queries like a video game is played, head passes in front of camera, and a delicious japanese dish describe broad visual concepts.

When the benchmark accepts one clip, semantically valid results can still land outside the target window. That points back to the dataset as well as retrieval.

Subjective or context-heavy captions are harder to anchor visually. Most beautiful resort i have ever seen depends on context and judgment.

A clip can match the text while missing the dataset's selected moment. For those queries, the label is not a complete representation of the caption.

Discussion

These examples make the benchmarks more useful when read carefully.

For retrieval tasks, a caption can describe an event, a repeated action, a broad visual category, or a subjective impression.

When that caption is paired with one accepted timestamp, the benchmark becomes sensitive to annotation coverage. If the timestamp is wrong or incomplete, the score can hide a valid retrieval.

That is worth keeping in mind when reading scores. Some misses are retrieval misses. Some are places where the dataset is wrong, even when the dataset is considered a standard benchmark.

Before evaluating a system against a benchmark, the dataset itself needs review. That can be manual review, LLM-assisted review, or both.

In video retrieval, one ground-truth window is not always the only correct answer.