Glossary
Speaker detection
Speaker detection is the part of video and audio analysis that works out who is speaking and when. Audio-side, it is often called speaker diarization: splitting a transcript into “speaker 1” and “speaker 2” segments. Video-side, active speaker detection matches the voice to a face on screen, usually from lip movement and face tracking. Clipping and reframing tools use it to decide where to point the vertical crop in a podcast or interview, and to label captions by speaker. It fails in predictable places: people talking over each other, off-camera voices, profile shots, masks, and wide shots where faces are small. That is why speaker-based crops need review.
In a HypeNest pass
In HypeNest, reframing aims to keep the active speaker in the vertical frame on interview and podcast clips. Watch the switches between speakers during review — a crop that lands on the listener while the guest delivers the punchline is a reason to recut or drop the clip.
FAQ
- What is the difference between speaker detection and diarization?
- Diarization is the audio part: which voice spoke when. Active speaker detection links that voice to a face in the picture so a crop can follow it.
- Does speaker detection work with more than two people?
- It can, but accuracy drops as faces get smaller and people interrupt each other. Panels and group calls need the most review.
Try the workflow
Draft captioned short clips from one long video, review every cut, then publish to YouTube and TikTok.
Start freeRelated terms
- AI clip generatorSoftware that proposes short vertical clips from a long video — often with captions — for human review before anything i…
- Talking headA talking-head video is one where the main picture is a person speaking to camera or to an interviewer — podcasts, inter…
- Auto-reframeAuto-reframe is automatic cropping that moves a vertical (or square) window across a landscape video so the subject stay…
