How to Find a Person in Your Videos Without Uploading Them
Video face search works by sampling frames, matching them against a reference photo, and merging hits into time ranges. Here is how to find the moment someone appears — with nothing leaving your computer.
Published 2026-09-24 4 min read
- video search
- privacy
- how it works
Finding a person in a photo is one problem. Finding the moment they appear in a two-hour video is a different one — and it is the moment most people give up and start scrubbing the timeline by hand.
The reason it feels hard is that video is not searchable the way text is. There is nothing to grep. The practical approach is the one video editors have used for years, automated: sample frames, look for faces in them, and turn the hits back into timestamps.
Why cloud video search is a bad trade here
Every consumer “search your video” feature works the same way: you upload the video, someone else’s GPU processes it, and you get timestamps back. For a holiday clip, that is convenient. For anything you would not want on someone else’s server — interviews, client footage, family video, anything with a minor in it — it is not.
Uploading is also slow. A few hours of footage is tens of gigabytes; the upload usually takes longer than the analysis. Doing it locally avoids both problems at once, and it is why face search on video belongs on your own machine.
How video face search actually works
SnapByFace treats video as a sequence of stills:
- Look at the video about once per second. Roughly one look per second is the sweet spot: often enough to catch someone walking through the shot, sparse enough that a long video does not take all day.
- Detect and encode faces in each sampled frame using the same bundled models used for photos.
- Compare each of those moments against your reference photo and keep the ones close enough.
- Merge consecutive hits into time ranges. Twenty matching frames in a row become one entry — “02:14–02:34” — rather than twenty separate results.
That last step is the one that matters for usability. A raw list of frame numbers is technically correct and practically useless; merged ranges tell you where to put the playhead.
:::note Video support covers MP4, MOV, AVI, MKV, M4V, WebM and TS. Photos cover JPG, JPEG, PNG, BMP, TIFF, WebP and HEIC/HEIF. Camera RAW files (CR2, NEF, ARW and similar) are not supported. Both go into the same index, so one reference photo finds the same person across photos and video at once. :::
The workflow
- Add a reference photo of the person — one clear, front-facing shot.
- Add the folders holding your video files. SnapByFace walks them recursively, so a messy years-old folder tree is fine.
- Start the index. It runs in the background, and you can pause, resume or cancel it; after an unexpected quit, it resumes from where it stopped.
- Open the results. Each video hit carries a time range — double-click to jump the preview straight to that second.
How long does it take?
The work scales with how long the video is, not how big the file is: an hour of footage means about 3,600 moments to look at.
Two things keep this practical on large collections:
- Incremental re-indexing. Adding new footage does not re-analyse what is already indexed.
- Speed at any size. Search stays quick even across hundreds of thousands of faces, so results arrive while you are still at the screen.
If you want to see how it behaves before committing your whole archive, index one folder first and check the top hits against your own expectations. That is the cheapest way to learn whether 0.45 is the right threshold for your footage.
What stays on your machine
Recognition and search run locally. No video, no frame, and nothing about the faces in them is uploaded: everything the app learns stays in ~/.snapbyface/. SnapByFace reads your media and never writes to it — no renaming, no moving, no re-encoding.
The only outbound data is an optional device statistics ping containing machine code, app version, operating system, language and activation state — no filenames, no footage, no search activity. It is on by default and can be turned off in Settings.
When the results will disappoint you
Worth knowing before you start:
- Fast cuts and heavy motion blur produce moments where the face is not clear enough. A brief appearance can also fall between two checks.
- Very small faces in frame (a wide shot of a crowd) are the hard case for any face search.
- Heavy appearance change, or a reference photo that looks nothing like the person in the footage, lowers the score below the threshold — try adding a second reference photo.
Try it on your footage
The trial runs the full local search and shows up to 5 matches per person and 20 per search, which is enough to tell you whether it finds what you are looking for in your own material. During the trial you can also start up to 100 analyses in total.