TWO-SPEAKER SPLIT SCREEN
Keep both podcast speakers visible in every vertical clip
Split-screen podcast clips solve the failure a moving single-person crop creates: the host disappears during a reaction, the guest is cut off, or the frame jumps on every turn. ClipFlap can keep both speakers in a stable stacked composition.
- 1
Detect the voices and visible people
Transcription can identify speaker turns while sampled face detection locates the people in the source frame. The signals are combined instead of treating every detected face as an active speaker.
- 2
Choose a composition per clip
A suitable two-person shot can become two stacked crops, while a single-speaker or unsuitable scene can retain another layout. The decision is made from the actual source framing.
- 3
Render captions with the conversation
Word-timed captions sit with the vertical composition and can use speaker-aware styling, keeping the dialogue readable without a constantly moving viewport.
Two-voice detection is more than counting faces
Podcast frames often include photographs, monitors, posters or a face printed on clothing. A useful detector must distinguish those from people participating in the conversation, then relate speaker turns to stable crop regions. Sampling and smoothing reduce the camera-chasing effect of reacting to every frame independently.
The goal is not to switch layouts on every sentence. It is to choose a composition that remains intelligible for the clip: both participants visible when the exchange needs them, or a focused single-person frame when that is the honest source structure.
Common crop failures the stacked view avoids
A narrow crop can remove a listener's reaction, cut a leaning speaker at the edge, or zoom toward a false face in the background. Constant left-right movement is also tiring in a rapid exchange. A stable stacked split keeps two known crop regions on screen and gives each person predictable space.
Not every shot can support a clean split. Burned-in lower thirds, remote-call windows and extreme camera positions may call for a wider treatment. Reviewing the rendered clip remains the final check before publishing.
REAL SPLIT-SCREEN CLIP
Watch the split-screen clip →How the framing decision changes
| Source situation | Risk with one moving crop | ClipFlap approach |
|---|---|---|
| Two people alternating | Repeated jumps between faces | Stable stacked crops when the shot supports them |
| Poster or monitor face | False crop target | Movement and size checks reduce false targets |
| Speaker near the edge | Face or gesture cut off | Crop planning preserves a safer subject region |
| Existing remote layout | Windows cropped unpredictably | Source composition is considered before reframing |
Split-screen podcast questions
Do I need two separate camera recordings?+
No. A single wide source that shows both people can provide two crop regions. Existing multi-camera edits and remote-call layouts can also be evaluated shot by shot.
Why not follow only the person who is speaking?+
A moving crop can hide reactions and jump constantly during quick exchanges. A stable two-person composition often reads better in a conversational clip.
What if a television or poster contains another face?+
Face detection is combined with size, position and movement checks to reduce false targets. The rendered output should still be reviewed before publication.
Are captions compatible with the stacked layout?+
Yes. Captions are timed to the spoken words and placed as part of the final vertical render so they remain readable alongside the two crops.
Can I manually force the layout for a clip?+
Layout override is not exposed in the current dashboard. ClipFlap chooses from the source and speaker signals, and you review the rendered result.
Related two-speaker workflows
Test a two-person episode
The free trial includes 10 minutes of processing and needs no card. Trial clips carry a discreet clipflap.com mark; paid plans remove it.