The direct answer

Highlight detection in a tool like ClipFlap reads the transcript of the episode, not the video. It looks for sentences strong enough to stand alone, builds a candidate clip around each one so that the idea starts and finishes inside the clip, discards candidates that match a list of things that never travel, and then scores what is left. Framing and captions come after selection and work from the picture and the sound. Knowing this explains both why the selection is good at finding a thought and why it can miss a moment that only happened on screen.

Step one: the transcript, with speakers and timing

Everything starts with a transcript that has a timestamp on every word and, when the source allows it, a speaker label. The transcript has to be in the language actually spoken; a Moroccan podcast transcribed into Standard Arabic gives the selector words the guest never said. Punctuation matters more than it seems: without sentence boundaries, a five-minute monologue is one sentence and there is nowhere honest to cut. When a transcript arrives unpunctuated, punctuation is restored before selection.

Step two: windows, so a long episode is read evenly

A three-hour episode is not read in one pass. It is split into overlapping windows of several minutes, and each window is asked for its strongest moments. This is what stops a quiet, excellent stretch late in the episode from being ignored because the first hour was loud. It also means candidates at window edges are seen twice, which is how duplicates get merged rather than shipped.

Step three: peaks, then boundaries

Within a window the selector first identifies the handful of sentences that would work as a quote: a claim, a reversal, a story's payoff, a number. Each becomes the anchor of a candidate. The candidate then grows backwards to the first sentence that stands alone and forwards to the point where the thought closes, with an explicit rule that a clip never ends on a connector, an unanswered question, an open list or an announcement not delivered. If the thought cannot be closed within a reasonable length, the peak is dropped in favour of another rather than truncated. This is the same method a human editor would use; the difference is that it is applied to every window with the same patience. The human version is described in how to find the best moments in a podcast.

Step four: the negative list

Some passages are loud and worthless as clips, and the selector is told so explicitly: intros, housekeeping, meta talk about the episode, generalities, passages that depend on something visual, and second clips about a peak already taken. Saying no is half the job, and early versions that lacked this list produced confident clips of "welcome back, today we have".

Step five: a score and a hook per clip

Each surviving candidate is scored on its own, not in a batch, on axes like how self-contained it is and how strong its opening line is, and gets a title card written to be faithful to what is said. Scoring one clip at a time turned out to matter: scoring a batch produced ranks that drifted between runs and cards that ended up on the wrong clip. The score is a sorting aid for review, not a prediction of performance, and we do not present it as one.

What happens after selection

Selection produces time ranges. Everything visual comes next and works from the picture: faces are sampled a few times a second, tracked over time, and the layout is chosen per camera shot: single, stacked for two people, wide when the camera outruns the crop, webcam-over-game for streams. Captions are timed word by word from the audio, and the hook card is placed where it covers no face. None of this feeds back into which moments were chosen.

What the selector cannot see

Because selection reads words, it is blind to the picture. In an internal test on a football game recording, the selector chose three moments from the commentary, and only one of them showed the goal on screen; the other two were the commentator describing something the game camera had already cut away from. On a podcast this is rarer, because the words are the content, but it still happens: a guest holding up an object, a chart being pointed at, a reaction that was the whole joke. The review step exists for this.

The selector also inherits every flaw in the transcript. A proper name misheard becomes a hook with the wrong name. A dialect rewritten into a standard register produces boundaries in the wrong places, because the sentence structure the selector reads is not the one that was spoken. Which is why transcript quality is treated as a gate: a transcript that looks unusable stops the job before selection rather than producing confident nonsense.

What this means for you

Trust the selection for what it is good at, which is finding complete thoughts across a long episode without fatigue, and review it for what it cannot do, which is see. Check that each candidate makes sense with no context, that the hook is faithful, and that nothing visual was the point. The seven checks are in what makes a good podcast clip, and the full workflow in the complete guide to podcast clips. To see the selection on your own episode, paste a link into the AI podcast clip generator.