Client systemCinematic search · Multimodal retrieval

ShotDeck

Multimodal search across millions of cinematic images.

Partner / contextShotDeck
Binding constraint

Preserve cinematography-aware meaning while ingestion, retrieval, infrastructure, and dependencies change at scale.

Binding constraintPreserve cinematography-aware meaning while ingestion, retrieval, infrastructure, and dependencies change at scale.
01Film + multimodal query
02Represent + select
03Searchable library
04Hybrid retrieval
05Ranked shots

Retrieval that understands framing, lighting and lens, not just pixels, so a filmmaker can describe a shot and get it back in under a second.

3M+Live production stills
200K+Frames sampled per feature
<10 min90–150 minute feature processing
Sub-secondRetrieval

Filmmakers search through framing, lighting, lens, composition, color, motion, and references that may arrive as text, a still, a clip, or spoken direction. The library also had to grow from full-length films.

Nixense engineered a shared retrieval layer for text, image, video, and voice. Semantic similarity carried conceptual relation; cinematic structure preserved framing, lighting, lens, composition, color, and motion.

The system sampled more than 200,000 frames from a 90–150 minute feature, scored and reduced them while preserving scene coverage, and indexed representative results. Workers expanded for processing and returned to an idle footprint afterward.

Evidence boundary

Processing time describes the recorded 90–150 minute feature workload; it is not a guarantee for every source or infrastructure configuration.

The auto-tagging system began at a 3% exact full-record match rate. A similarity-curated set of 20,000 examples outperformed 100,000 random examples by more than twenty points and supported a 71% exact-match result in the recorded benchmark.

The lesson was not that smaller data is always better. The training set had to represent the decision boundary the product needed.

When an external capability reached sunset, Nixense rebuilt embeddings, color, and tagging in-house, reprocessed the corpus, and kept the transition away from the product surface.

Routine orchestration expanded from one idle worker to roughly 40–50 during processing and returned to one. A separate always-on comparison moved from approximately $3.20/hour to $0.80/hour.

The hourly figures describe a separate always-on comparison, not the rate for the dynamically scaled workers.

Record close
The result was not one better model. It was a complete path from source film to searchable visual record, and from creative intent to a relevant result — built to keep operating as the corpus grew.

Working on a problem like this?

Tell us what the system needs to do, what exists today, and where you need help.

Discuss a system
Principal review[email protected]Pakistan · Working internationally