copydnn

queriesstreamsrefsclustersmatchescrawlerapi
settings

bleeding-edge content retrieval · dual-signal video, image & audio

Find where a video came from.

copydnn matches any video, live or recorded, against your catalogue and names the source and the exact seconds it came from, however it has been altered: cropped, mirrored, sped up, or buried in a corner. It reads two signals at once — the video fingerprint and the audio track — so a copy that survives one is still caught on the other. Production-ready, with a very low false-positive rate even on live sport.

Drop files or paste a YouTube link, or choose files

files run one after another; a YouTube link is fetched once and only its fingerprint is kept. Queries land in Queries, references in References.

Nothing to test with? Browse the live library of real copyrighted footage →

64msmedian query
<0.1%false positives
10.000h+on one machine
0.3sfrom ground truth
robust matching

Finds the copy through every edit

The problem is old and hard to solve. Once a video leaves your hands it stops being your file. It comes back re-encoded, trimmed, cropped, mirrored to dodge takedowns, recoloured, or tucked into the corner of somebody else's broadcast. A checksum survives none of that, since a single changed pixel is enough to break it, and filenames and metadata were never proof to begin with.

copydnn works from what the frames actually show, so the copy turns up despite all of that handling. Every claim it makes arrives with the reference it matched, the exact seconds it came from, and the frames that matched.

On the right is one shot from the demo footage, put through the attacks our benchmark uses. Every version still matches, and every tile is a live query. Click one and that exact clip runs against this library.

the original clip
the originalrun it →
the same clip mirrored
mirrored✓ matches
the same clip cropped to 35 percent
cropped to 35%✓ matches
the same clip recoloured
recoloured✓ matches
the same clip rotated 15 degrees
rotated 15°✓ matches
the same clip sped up by half
sped up 50%✓ matches
live · the showroom

Try it on real footage

Each clip above is the same shot of Big Buck Bunny, altered on purpose to prove copydnn still finds it. A handful of clips can only prove so much, so we built a crawler to run the same detection on real footage at scale. It follows the official channels of studios, networks, and labels and fingerprints every trailer and scene the day it posts, so the library keeps filling with genuine copyrighted footage and stays current on its own. The video itself never lands here. copydnn keeps a few descriptors per second and plays the original from YouTube in place. Browse it, query against it, and watch a stray re-upload snap back to its official source.

  1. 1
    Real material, kept current.

    Official channels, fingerprinted the day they post. This week's trailers are already indexed.

  2. 2
    Fingerprints, not footage.

    The frames stay on YouTube. copydnn keeps a few descriptors per second and plays the original in place.

  3. 3
    Open to explore.

    Browse the library, open the duplicate map, or run your own clip against thousands of hours of it.

watching now
Caught recently see all matches →
Browse the showroom → how the fingerprint works
built for live sport

Where every frame looks the same

Reliably detecting live sport is some of the hardest content there is, because the visual similarity runs so extreme. The same backgrounds, the same camera angles, the same broadcast overlays repeat from one event to the next. Two separate matches can run 95% identical frame to frame, the exact case where a weaker detector raises a false alarm. copydnn is tuned for it. It fires only on a sustained, distinctive run through time, so a real re-upload snaps back to its source while two separate matches stay firmly apart. On sport, that keeps the false-positive rate near zero.

  1. 1
    Everything looks alike.

    Same court, same cage, same angles and graphics. Two separate events overlap almost completely, the hardest case a matcher meets.

  2. 2
    False positives are the real risk.

    A weaker detector treats that overlap as a copy and flags two separate matches as the same footage. On live sport it would happen all day.

  3. 3
    copydnn matches on evidence over time.

    A distinctive run, temporal aggregation, and a robust fit are what confirm a match, so a real re-upload is caught while two separate events stay apart.

A Wimbledon singles match on Centre Court, from the broadcast camera
Wimbledon · one match
≈same court
A different Wimbledon match on the same Centre Court, same camera
Wimbledon · another match
per-frame hits over the cliplongest aligned run 0.2s < 3s needed
Same court, different match · kept apart
A UFC octagon during one bout, shot from the broadcast camera
UFC · one bout
≈same cage
The same octagon and camera during another bout on the same card
UFC · another bout
per-frame hits over the cliplongest aligned run 0.2s < 3s needed
Same cage, different fight · kept apart

Low-res stills from Wimbledon and UFC broadcasts, shown for illustration.

who it's for

From rights holders to in-house screening

Index your footage once and the same fingerprints answer very different questions. Protect a catalogue, screen uploads, trace a fragment, clear out duplicates, or keep everything behind your own firewall: it is one library doing all of the work.

rights holders

Know where your footage went

A catalogue that took years to produce reappears on other channels a few seconds at a time, and rarely untouched. Index your library once and every suspect clip can be checked against all of it. When something matches, you get the reference, the exact seconds it came from in that reference, and a result page you can hand to whoever needs convincing.

your catalogue3:30matched

platforms

Screen uploads before they go live

Every incoming video is one call against the library, settled in milliseconds. A clear match can be blocked or flagged on the spot, a possible match can go to a human with the match page attached, and everything else passes through without anyone noticing the check happened.

uploads?publishreviewreject

newsrooms and archives

Trace a fragment to its source

Someone hands you thirty seconds of footage, or a single screenshot, and asks where it came from. If the source is in your library, copydnn names the recording and the moment, even when the fragment was mirrored or re-timed along the way. Images find their place inside videos, and clips find indexed stills.

one still3:35bunny.mp4

collections

Find what you already have twice

Archives grow duplicates the way attics do, gathering the same footage under different names, encodes and resolutions. Because every reference can be queried against all the others straight from its stored fingerprint, finding them is cheap, and it works at the scale of thousands of hours on a single machine. New inserts are checked the same way, so a copy is refused at the door and the refusal names what it duplicates.

same footage

in-house content screening

Host it inside your own walls

Large platforms run screening in-house, where it can reject copyrighted material or known terrorist footage at upload time without any third-party service ever seeing the traffic. copydnn is built for exactly that. The index, the fingerprints and the results all live on your machines, behind your firewall, and the library grows to the ten-thousand-hour tier on ordinary hardware.

your infrastructureindexuploadsreject

inside the engine

How does it work?

A query runs in four stages, and each one narrows the work waiting for the next. Every stage below comes with a real example you can run as you read.

01 · learned global embeddings

Every frame becomes a point in space

A neural network reads each frame and writes it down as 512 numbers. The technique is called a learned global embedding, and ours is SSCD, trained on millions of examples of how copies get altered out in the wild. All that training makes the fingerprint hard to fool. Mirror a frame, crop it, recolour it or push it through another round of compression, and its numbers barely shift, while a frame of something else lands somewhere else entirely.

The translation runs one way only, so the picture cannot be read back from a fingerprint. Near-duplicate moments are stored only once, and that is how an hour of video comes to about eighteen thousand entries rather than a hundred thousand.

try it:

A frame from the demo footage becoming its numbers, 48 of the 512 shown, tinted by size. Run the same frame through any of the disguises above and these numbers barely change.

02 · candidate retrieval

All those numbers go into one big index

Every fingerprint we store goes into a large, purpose-built database that answers one question very quickly. Which stored moments have numbers closest to these? A smaller library gets there by checking every entry outright. Once a collection reaches archive scale the work moves to FAISS, the vector search library Meta built for billion-entry workloads, which keeps thousands of hours searchable on a single machine.

This stage stays deliberately generous, because a moment lost here is lost for good. Each query frame gathers dozens of nearby candidates while they are still cheap and hands the whole lot onwards. Sorting the real ones from the coincidences belongs to the next stage.

push it harder:

A query frame's point, its nearest stored neighbours ringed. Everything else in the index is never touched.

02 · at archive scale

Ten thousand hours on one machine

Retrieval stays exact for as long as it can afford to be. Up to a few hundred thousand stored moments it compares every vector directly, which covers around 244 hours of video in a quarter of a second.

Past that point the FAISS index above takes over. It carves the embedding space into cells, the way the diagram shows, and a query visits only the few cells nearest to it, so the cost of a lookup follows the number of cells probed rather than the size of the archive. Each vector is also compressed down to a 64-byte code, thirty times smaller than the original, which keeps memory use flat at roughly 1 GB per 1,000 hours of video.

The compact codes decide which rows compete, not which one wins. They pick a shortlist, and that shortlist is scored again against the exact vectors kept on disk, so the ranking you read is the exact one. Throughput stays high at both ends. Ingest runs at about fifty times realtime on a laptop GPU, and a fingerprint query comes back in nine milliseconds, which works out at more than a hundred queries a second on a single thread.

The IVF idea, drawn to shape. The space is carved into cells, and a query visits only the nearest few, shown tinted. Everything else stays cold on disk at 64 bytes per vector.

03 · temporal aggregation

Time separates copies from lookalikes

Two frames that look alike are only a hint, since the world is full of footage that resembles other footage. What makes a copy a copy is that it travels through time alongside its source. So every candidate gets plotted, when it happens in your clip against when it happens in the reference. A real copy draws a straight ridge across that plot, because the two timelines advance together. A chance resemblance just scatters.

This is the stage that keeps false alarms out. Evidence has to hold up in time, frame after frame, across seconds of video, before it counts for anything at all. On our benchmark that discipline is the whole distance between zero false positives and educated guessing.

try to fool it:

hover the surface to read any pair
a real correlation surface: every query frame against every reference frame

04 · robust fitting, RANSAC

A line that outliers cannot drag

That ridge still has to turn into a precise line, and averaging would be a poor way to get one, since a single stray pair drags an average wherever it likes. So the pairs hold an election instead, using a classic technique called RANSAC. Random pairs each propose a line, reference = rate · query + offset, and whichever proposal the most pairs agree with wins. The agreeing pairs carry the fit, and the scattered ones are simply outvoted.

The slope of the winning line turns out to be the playback speed itself. A copy running fast gets modelled exactly and reported as rate 1.5, with the offset naming the second the copy begins.

re-time it:

Every dot is a candidate pair. The ones that line up elect the line and carry the fit, and the scattered ones are outvoted.

05 · the verdict

Reading what comes back

What comes out at the end is a claim you can check for yourself. This stretch of your clip is that stretch of that reference, at this speed. Every result page carries it as two short lines. Here are both, taken from the sped-up example above, with an explanation of each number underneath.

1STRONG bunny.mp4 2best 0.994
3query 0.0–9.0s → reference 3:30–3:43 4rate 1.454× 510 frames 6mean 0.953 7resid 0.12s
1strong · possible

strong asks for a run long enough, similar enough and dominant enough to account for the whole query rather than a fragment of it. possible covers a match that holds together while falling short of that bar. Both thresholds come from measuring what content we know to be absent actually scores.

STRONG possible
2best

The single most similar pair of frames, on a scale from 0 to 1. Taken alone it is an anecdote, and the run behind it is what turns it into proof.

0 1 0.994
3the run

One continuous appearance. Seconds 0 to 9 of the query are 3:30 to 3:43 of bunny.mp4. A single query can hold several runs, and one reference can turn up in several places along it.

query bunny.mp4 · 3:30–3:43
4rate

The slope of the fitted line, measured in reference seconds per query second. 1.000 means the copy runs at source speed. The 1.5× clip fits at 1.454, since sampling a couple of frames a second resolves the slope only so finely.

1.454× 1.0×
5frames

How many sampled frames sit on the line. Any two points at all will make a line, so the weight of the claim lives in this number. Ten independent frames agreeing on one mapping have stopped being a coincidence.

6mean

Average per-frame similarity along the run. Cropping and re-encoding drag this down while leaving the alignment intact, so a middling mean backed by plenty of frames still counts as strong.

0.953
7resid

How tightly the frames hug the line, as the median gap in seconds between where each frame landed and where the line says it should be.

±0.12s
looping

The same few stretches of a reference recurring right through the query, which happens constantly in rolling broadcasts. The content really is there, repeated rather than continuous, and the result says as much in plain terms.

the same stretch, seen 4×
see one for yourself

or open a live result →

Point it at the library itself

A query settles one question at a time. Ask every reference that same question about every other one, and what comes back beats a list. You get the shape of your collection, showing you which files are the same work, which ones only share a little footage, and which ones stand on their own.

clusters · the library as a graph

A map your library draws for you

Every match that survives becomes a link, and the references it joins land in the same cluster. Those links are ordinary query results held to the same standard as any query, so a cluster remains something you can open and check one link at a time. Each rests on what the frames actually show, exactly as a single query does.

The map keeps two things apart that a duplicate finder usually runs together. Same work means a whole copy of one thing under another encoding, crop or name. Shared footage means a stretch that turns up inside otherwise unrelated videos, such as a title sequence, a news clip or a stock shot. The distinction matters the moment a series appears, since merging on a shared opening would collapse a whole season into one blob. Every link says which of the two it is, and you set the bar for both.

The whole thing runs on the fingerprints you already have. Your media stays exactly where it is, and a rescan re-reads what ingest wrote the first time around, so redrawing the map costs you nothing but a moment.

see it on real data: open the cluster map →

Links look the same here as they do on the real map. Colour runs from cyan to amber with match strength, solid for a whole copy, dashed for footage two videos happen to share, faint where a match is only possible. In a healthy library plenty of files stand alone, and seeing that is useful too. On a 3.817-reference collection the map draws 291 clusters over 1.047 links and leaves the other three quarters standing alone.

Want to know more?

The five chapters above cover the whole pipeline. Underneath them sit a dozen smaller decisions about which frames to look at, which parts of each frame, how much any one similarity is worth, and how it all stays fast once the library gets big. Each one below is closed until you open it.

01 How a frame becomes 512 numbers the descriptor itself

The model is SSCD, a network Meta trained specifically for copy detection. Most vision models learn to say what is this a picture of. This one learned something narrower and more useful here. It saw millions of examples of how copies get altered in the wild, from crops and mirrors to recolouring, re-encoding and pasted-on logos, and it was rewarded for giving the altered version the same output as the original.

A frame goes in at 224 × 224 pixels. A stack of convolutions reduces it to a grid of features, and a pooling step called GeM squeezes that grid into a single list of 512 numbers. Scaling the list to unit length is what lets two frames be compared with one multiply-and-add, the cosine similarity behind every score on this site.

One detail matters more than the rest. SSCD's training includes a term that actively spreads unrelated descriptors apart, so the gap between "this is a copy" and "this just looks alike" stays wide, and stays roughly the same width everywhere in the space. That consistency is what lets one global threshold work across an entire library instead of being retuned per video.

The numbers travel one way only. Nothing reconstructs the picture from them, so a fingerprint gives away nothing about what it describes.

Pixels in, one point on a sphere out. Unit length is why comparing two frames is a single dot product.

02 Describing pieces of a frame, and the whole the region ladder, and why it is lopsided

One descriptor per frame is a summary, and summaries dilute. Someone takes a small crop of your video and drops it into the corner of theirs. The frame they publish is mostly their content, so its 512 numbers describe their frame, and the match to yours is lost.

So each frame is described several times over, once for the whole picture and once for each of a few windows inside it. A partial copy can then match window to window rather than needing the whole frames to agree. The windows are concentric because the edits that dilute a global descriptor tend to be concentric too: a centre crop, a letterbox, a picture-in-picture sitting in the middle.

The two sides get deliberately uneven treatment. A reference is stored forever and every window multiplies the size of the index, so a reference gets three, permanently. A query is temporary, its cost gone the moment it is answered, so a query can afford far more: switch on every window capability and it is described fourteen ways, each also embedded mirrored, for twenty-eight views. That asymmetry buys the recall of a fine-grained search at the storage cost of a coarse one. Because it is transient, it is yours to spend or not, per search (chapter 07). The default is the whole frame alone.

An earlier version also cut each frame into quadrants. It cost twice as much and found everything the nested windows already found, so we dropped it. There is a real ceiling here as well. Every extra window is another chance for something unrelated to look similar, which is why the grid stops where it does.

Three windows kept per stored frame; up to twenty-eight searched per query frame when every capability is on. The cost of the wide side is paid once and thrown away.

03 Letting the clip say where its content is adaptive region detection

Nested windows are a guess about where an edit put the content, and real edits are under no obligation to cooperate. Put two videos side by side and the split runs down the middle, where no centre crop reaches. A centre crop of that frame captures half of each picture and resembles both sources equally badly. Letterbox a clip and the whole-frame descriptor spends much of itself describing black bars.

So before searching, the query gets measured. Two things give an edit away.

Flat borders. Scanning inward from each edge, rows and columns with almost no variation are padding: black bars, a plain backdrop, the surround of an inset. What survives the scan is the real content box. This single test catches letterboxing, pillarboxing, and a picture-in-picture sitting anywhere in the frame, including the corners that concentric windows can never reach.

Seams. A composite has a boundary. Sometimes that boundary is a visible divider, which shows up as a narrow band of flat columns. Sometimes the two pictures simply abut, and the giveaway is neighbouring columns that stop resembling each other far more abruptly than ordinary footage ever manages.

A layout only counts once it recurs across the clip. Editing persists for a whole shot at minimum, while a moment that merely looks like an edit passes. That distinction is what keeps a cut to black or a dark sky over a dark sea from being read as letterboxing. It also keeps every batch the same shape, which is what lets all the windows of all the frames go through the network in one pass.

Found windows join the ladder, capped at four, and anything that duplicates a window already there gets dropped. The result raises recall and lowers false positives at the same time, because a region that really is one picture makes a better match than a guessed one straddling two.

The profile under the frame shows how much each column varies. Flat runs at the edges are padding, and the cliff in the middle is a seam.

04 Stepping past blank frames content-aware sampling

Watching every frame of a video would mean thousands of descriptions per minute for footage that barely changes, so the sampler takes a couple of frames per second and moves on. That works well, and it carries one weakness worth being honest about. A fixed grid is predictable, and anything predictable can be aimed at.

Blacking out a single frame every second costs one line in a video editor and passes a viewer completely, since it is one frame in twenty-five and gone before the eye registers it. It lands squarely on a regular sampling grid. A sampler that keeps whatever the clock hands it then spends a large share of the clip describing black rectangles, and those descriptions carry no identity, so the signal for a real match quietly thins out. This is a published attack with measured results behind it.

The counter is small. When the frame the clock points at carries nothing, meaning near-uniform pixels of any colour, the sampler steps sideways and takes the next frame that carries something. The timestamp moves by a few hundredths of a second, well inside the tolerance the timeline fit already allows, and what gets described is a frame that means something.

Stepping is bounded, and the bound is the interesting part. When a clip genuinely fades to black for several seconds, skipping past all of it would leave a hole in the timeline exactly where the query really was, and a fit that sees an unexplained gap is worse off than one that sees darkness. So after a few tries in a row the sampler takes what it is given. On clean video it behaves exactly like the fixed grid.

Ticks are where the clock wanted to look. Where that frame is blank, the sample steps to its neighbour.

05 Not all similarity is worth the same score normalisation

A similarity of 0.6 sounds like a fact, and it hides one. Some frames sit close to everything: a title card, a fade, a stretch of anonymous studio footage, a talking head on a plain backdrop. Their 0.6 means very little. A frame with unusual composition scoring 0.6 against one reference and nothing else means a great deal. A single global threshold reads both as the same number, and thresholds are how verdicts get made here.

So every query frame answers a second question: how well do you score against the collection in general? A small, fixed sample of the library stands in for material at large, and a frame's similarity to its nearest few of those becomes its background level, a measure of how generic it is. Discounting that level from its scores leaves the part that actually distinguishes one reference from another, and that is what the match gets weighed on.

Two decisions keep this honest. The discount applies to the query side only. Textbook recipes discount both sides, which works when the background sample comes from somewhere unrelated to the collection being searched. This library exists to hold things that resemble each other, so on the stored side the same arithmetic penalises a reference for resembling its own near-duplicates, which is precisely backwards.

The discount is also centred, so an ordinary frame stays where it was while an unusually generic one loses ground and an unusually distinctive one gains. Every score therefore stays on the familiar scale you read on a result page, instead of the whole system sliding down by some arbitrary amount. The effect is higher recall at no cost in false positives.

The pale part of each bar is the background level, meaning what that frame scores against material at large. What is left over is what counts.

06 Storing each shot once novelty-based deduplication

A locked-off shot of someone talking for thirty seconds gives a descriptor the same 512 numbers hundreds of times over. Storing all of them inflates the index and adds nothing anyone could ever find.

So a sampled frame earns its place only once it has drifted far enough from the last one kept, meaning its similarity to that stored frame has fallen below a threshold. A static shot collapses to a single entry however long it runs. A fast-cut sequence keeps almost everything, because almost everything in it is new. The rate follows the footage rather than a number someone guessed.

What makes this pleasant is that it needs no scene-detection step. The descriptors already know when the picture changed, since measuring change is exactly what they do. Shot boundaries fall out of the same comparison that does everything else.

One guard keeps it safe. However still a shot is, a frame goes in at least every couple of seconds. Timeline alignment needs timestamps spread along the video, and a shot that produced a single entry in four minutes would leave the fit nothing to anchor to.

Think of this as a search optimisation first and a storage one second. Every row left out is a row never compared, on every query, forever.

Sampled above, stored below. Still shots collapse into one entry while cuts survive, so the keep-rate follows the footage.

07 You choose what a query looks for one pass, four capabilities

Most queries are easy. A re-upload, a re-encode, a resize, a speed change or a recolour leaves the whole frame still looking like the whole frame, and one descriptor per frame settles it. That is the baseline every search runs, and on its own it runs in about a second per minute of footage.

The hard cases are the ones where the copy occupies only part of the picture: a deep crop, a corner inset, a split screen, a phone recording of a screen across the room. Those need the frame described in windows rather than whole, and each family of window is a capability you switch on: crops adds four concentric windows, insets a nine-cell grid, splits whatever the pixels' own geometry argues for, mirrored a flip of everything selected.

The cost is honest and stated on each one, because it is multiplicative: every window is another pass of the neural network over every sampled frame. All four together is roughly thirty descriptors per frame instead of one.

So the trade is yours rather than the engine's guess. Screening a stream of mostly-clean uploads wants the baseline; hunting a known leak through re-cuts and camrips wants everything, and the vectors are cached per query so the second look never re-embeds what the first already did.

One frame, described as many ways as you ask for. The baseline is the whole frame; each capability adds its own windows on top.

08 The ways a timeline fit can lie guards on the fit

Fitting a line through matching moments is the core idea, and left alone it is far too eager. Give it enough candidate pairs and some line always fits something. Most of the care in this system goes into refusing the fits that mean nothing.

A line has to touch real reference material. Mapping a long query onto two seconds of a reference satisfies a line beautifully while involving almost none of the reference. A run must land on several distinct reference moments to count.

One vote per query frame. A single frame can produce dozens of candidates, and inside a static shot most of them are near-duplicate reference moments scattered across the tolerance band. Counting all of them lets one frame outvote a hundred others and tilts the line. Each query frame casts one vote, its best.

Speed has a supported range, and its edges are a tell. When a fit reports a speed sitting exactly on the edge of what is supported, the constraint chose that number rather than the data. The honest value lay outside the range, which puts the content outside what counts as a copy here, so those fits get rejected instead of reported at the boundary.

Some agreement carries no information. Two videos that share a black tail or a standard title card agree perfectly and constantly, and a line drawn through that field passes every other test. The giveaway shows up across the library, where content-free frames match many different references equally well. Every run therefore carries a measure of how much of its support actually points at one moment, and runs built on ambience are held back from a strong verdict.

Repeats are one fact, not twenty. A rolling broadcast that loops the same segment all day produces many runs against the same stretch of reference at different query times. Those get grouped and reported as a single recurring match with its period, in place of twenty near-identical rows.

Three shapes on the same plane: a real run, a fit that squeezes the query onto one reference moment, and a band of agreement that points nowhere in particular.

09 Where the vectors live, and how they shrink storage, quantization and the index tiers

Every stored row is one moment of one window of one file. The descriptor leaves the network as 512 full-precision numbers, which comes to 2 KB per row, and that is more than the job requires.

On disk, half precision. Rows go down as 16-bit floats, halving the fingerprint. The values are already unit vectors in a small range, and all that follows is a dot product, so the loss stays far below anything that could change which row wins. The precision the network runs at is a separate question entirely, and there the same shortcut is genuinely destructive.

Small libraries: no index at all. Below a few hundred thousand rows, a query multiplies against every stored row on the GPU. Brute force is embarrassingly parallel and perfectly exact. It also stays affordable much further up than you might expect here, because this design keeps roughly three vectors per frame where a keypoint-based system would keep a thousand.

Mid-size: a proximity graph. Once a full scan stops being interactive, an HNSW graph takes over. Think of a navigable network where search walks from node to neighbouring node, always downhill toward the query. The settings deliberately trade a sliver of accuracy for build speed. A moderate number of links per node (M = 16) builds several times faster than a denser graph and gives up a fraction of a percent of recall, and build cost is what dominates the operational story. The search-time effort knob (ef = 256) stays generous, since that one only costs you when you use it.

Archive scale: quantization. Past a few hundred thousand rows, search moves to FAISS IVF-PQ, which does two things at once. IVF carves the space into cells, roughly four times the square root of the row count, and a query visits only a fraction of them, so lookup cost follows cells probed rather than archive size. PQ, product quantization, chops each vector into 64 short pieces and replaces every piece with the nearest entry from a small learned codebook, storing a one-byte code in its place. A 2 KB vector becomes 64 bytes, about thirty times smaller, and that is what turns thousands of hours into something that fits in memory.

Compression only ever shortlists. Compressed codes are good at saying which rows deserve a look and poor at ranking them, so they hand over a shortlist several times longer than the result. That shortlist gets re-scored against the exact vectors, memory-mapped from disk so only the shortlisted rows are ever touched. The ranking you read is the exact ranking, and the approximation merely decided who got to compete for it.

One structural oddity is worth naming. FAISS runs in a separate process, because it and the neural network library each bring their own parallel-computing runtime and loading both into one process is unstable. The split turns out to suit the work anyway. Describing frames is GPU work, searching is CPU work, and the two exchange nothing but arrays of numbers.

Three tiers, chosen by size. Bytes per vector fall as the collection grows, and the last tier hands its shortlist back to exact vectors before ranking.

10 Fast arithmetic, in the right places mixed-precision inference

Modern GPUs run half-precision arithmetic far faster than full precision, which makes describing frames in 16-bit an obvious win. Taking that win the obvious way destroys this model completely.

The pooling step is why. GeM raises every activation to a power before averaging, and in 16-bit those raised values run off the top of what the format can represent. Everything saturates, and every descriptor collapses onto the same vector. What makes the failure vicious is how quiet it is. Nothing crashes, no warning appears, the vectors still have unit length, and unrelated frames simply start scoring a perfect match against each other.

Autocast solves this by applying a policy per operation. Convolutions and matrix multiplies carry the bulk of the work and tolerate low precision happily, so they run in 16-bit. Reductions stay in full precision, including the pooling step that would otherwise break. The dangerous operation never sees the narrow format at all, and the descriptors come out matching their full-precision counterparts to five decimal places.

The gain shows up mainly in queries rather than ingest, for a reason worth knowing. Pulling frames out of a video file takes longer than describing them, so ingest runs at the speed of the video decoder. A query describes many windows per decoded frame, and that is where faster arithmetic earns its keep.

A start-up self-check embeds a handful of structured noise images and fails loudly if unrelated ones score alike. Should the collapse ever return by some other route, it will announce itself.

Sixteen-bit through the heavy layers, full precision through the one that would saturate. The lower strip is what a blanket conversion does.

the practical part

That's how it works. Here's how you use it.

The same engine, as an API

copydnn was built API-first. The page you are reading talks to the very same HTTP endpoints you would reach for, covering queries, inserts and the match views alike, so a script can repeat anything you just did by hand. Since the wire format is the fingerprint itself, your media can stay where it is throughout.

copydnn · one call to /v1 per line

        
The three verbs anyone starts with, and what comes back — each is one HTTP call: insert → POST /v1/references · query → POST /v1/queries · stream → POST /v1/streams.

integrate · 01

Point the CLI at an instance

The copydnn CLI ships with the engine. Give it --api, or set COPYDNN_API once, and the same verbs travel over HTTP instead of running locally. You type the same commands either way, and only the library moves.

# from the repo
pip install -e ".[web]"

export COPYDNN_API=http://127.0.0.1:5002
copydnn list
     ref  kind       dur  keyframes      MB  name
vid:1     video     9:56       1012    3.11  bunny.mp4
vid:2     video     0:14          7    0.02  neg_bars.mp4

integrate · 02

Insert: fingerprint where the media lives

One command turns a video into a .cdfp fingerprint on the machine that already holds the file. Adding it then pushes those few kilobytes instead of the gigabytes behind them. Where fingerprinting locally is impractical, every endpoint also takes a link to fetch or a plain upload.

# local: no network, media never moves
copydnn fingerprint film.mp4
  film.mp4.cdfp  63.9 KB  20 keyframes

# remote: pushes the sidecar, searchable on return
copydnn insert film.mp4.cdfp
video:3  film.mp4  20 keyframes  0.06 MB  ready

integrate · 03

Query: a verdict in one call

Query with a fingerprint or with the media itself. Back comes the claim you just learned to read, together with a permanent url where this UI renders the full match, from the synced players to the surface. It returns in milliseconds, so screening uploads at the door is exactly this one call.

copydnn query clip.mp4.cdfp
clip.cdfp: 20 frames sampled, 51 ms (tier fingerprint)
  STRONG   bunny.mp4  best 0.945
      query 0.0-13.7 -> reference 59.0-73.7  rate 1.037x

# the same as raw HTTP: JSON with a result url
curl -X POST -H "Content-Type: application/x-cdfp" \
  --data-binary @clip.mp4.cdfp $COPYDNN_API/v1/queries
{"verdict":"strong", "url":"/queries/42f9616d76c6", …}

integrate · 04

Compare: two files, no library

Sometimes there is no library, just two files and one question. Compare replies in the same shape and stores nothing, which makes it handy in a pipeline that only ever asks whether two files are the same footage.

copydnn compare clip.cdfp film.cdfp
STRONG  best 0.945
    query 0.0-13.7 -> reference 59.0-73.7  rate 1.037x  20 frames

Behind all of this sits an API reference covering each endpoint, each request shape and each response field one by one, with real captured requests and responses throughout.

Read the API reference →

Reference library

Insert references drop here, or choose

Every query is matched against these references. Open one to inspect it, query the library with it, or remove it.

Clusters

Recent matches

The crawler fingerprints official channels as they post. When a new upload reuses footage already in the library, it shows up here, newest first. Each row is one caught copy: who copied whom, how much of the video overlaps, and which of the two was uploaded first.

Queries

Add files drop here, or choose

Nothing here is a reference. These are the files you are checking, watchable and queryable without ever touching the index. Each one ends in exactly one of three ways. Run it against the library (as often as you like, the card keeps its latest verdict), promote it to a reference, or delete it.

Streams

Start new stream

A source watched while it is still playing. Frames are sampled, embedded and searched against the library on a window that keeps moving, so a match is reported while it is on air rather than after the recording is filed. What accumulates is a fingerprint, so ending a stream costs nothing: it becomes a query in Queries, or a reference in Refs. A reference is searchable from its first frames, without waiting for the broadcast to end. Feed one from the CLI (copydnn stream <url>) or start one here.

← streams

Start a stream

This instance opens the source itself and decodes it, so it has to be somewhere the engine can reach. If it is not, push frames to it instead: copydnn --api <host> stream <url>, or pipe ffmpeg into copydnn stream -.

What it is for
Sampling
Cancel
← streams

Stream

now sampling

the frame currently being searched

session

what has been on air

verdictreferencein the stream in the referenceratebest

the mapping, frame by frame

Settings

Bulk query

Run many queries in one go, under any search profile. Re-running the no-match ones with a heavier profile is how you hunt for false negatives, and every card keeps its latest verdict.

Bulk promote

Promote every no-match query into the reference library. That is the triage outcome where a file matched nothing and turned out to be new content worth protecting. Each promotion is a full ingest with dedup, so this takes ingest time per item.

Bulk delete

Clear one triage outcome — for example drop the matched ones once they are confirmed copies. Deleting everything regardless of state lives in Admin's danger zone.

Appearance

The theme follows the system until toggled; the choice sticks in this browser.

Maintenance

Re-ingesting recomputes every stored fingerprint under the current settings — what a REINGEST-tagged change needs to fully apply. It decodes and embeds the whole library again, and on a large library that is hours, not minutes. The request runs synchronously; leave the tab open.

Resetting settings drops every override on the Query and References tabs and restores the shipped defaults. Fingerprints already stored are untouched.

Danger zone

The data root holds three separate stores, query files, references, and stored searches, and clearing one says nothing about the others. Each button empties exactly one, and none can be undone. Emptying all three is what "clear the database" would mean.