Building the video backend for agents.

Field notes, build stories, research, and working projects from video infrastructure, retrieval, streaming, and agent tooling.

01 /
Highlights

Selected work
from across VideoDB Labs.

Research, field notes, and build stories worth a closer look.

Read JEPA: From Language Models to World Models

Research · Jul 7, 2026

JEPA: From Language Models to World Models

Why JEPA’s latent-prediction objective may shift AI systems from token prediction toward predictive world models for VLMs, VLAs, and embodied agents.

Read What video retrieval benchmarks taught us about ground truth

Field note · Jul 24, 2026

What video retrieval benchmarks taught us about ground truth

Manual review of MSRVTT, MSVD, VATEX, DiDeMo, and QVHighlights shows cases where the benchmark ground truth is too narrow, shifted, or ambiguous, so valid retrieved clips are scored as misses.

02 /
Field Notes

Recent production lessons
and technical notes.

Browse short notes on infrastructure, reliability, latency, retrieval, and the small fixes that matter in production.

03 /
Build Notes

Build notes
from working VideoDB projects.

Follow the architecture choices, API decisions, and tradeoffs behind apps, demos, and agent tools built with VideoDB.

04 /
Newsletter

Newsletters
from the VideoDB engineering team.

Read concise updates, technical notes, and build context from the team working on video infrastructure and agents.

05 /
Research

Research at the edge
of video and agents.

Papers, evaluations, talks, and notes on video understanding, retrieval, multimodal models, and agent systems.

Read Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

arXiv preprint arXiv:2604.11177 · 2026/4/13

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

Shivam Sharma, Sankalp Nagaonkar, Ashish Choithani, Ashutosh Trivedi

Benchmarks how internal reasoning traces affect video scene understanding in Gemini models, including where quality gains plateau and how tight budgets increase compression-step hallucination.

Read on arXiv