effects and edits.md
--------------------

 This looks like it's from one of the "AI faceless content" playbooks that circulate for building TikTok/Reels/Shorts channels around AI-generated visuals and voiceovers. Going through your questions in order, then breaking down the two secrets.

**Best editing techniques for images, video, and animation**

*Video*
- **Dynamic captions** — word-by-word animated text timed to speech. Probably the single highest-impact thing you can add, since most people watch muted.
- **Jump cuts** to kill dead air and keep pace tight
- **J-cuts/L-cuts** — next clip's audio starts before its video appears (or vice versa) so cuts feel smoother instead of jarring
- **Speed ramping** — slow into a beat or punchline, snap back to normal
- **Zoom punches** — a quick digital zoom on a key word or moment for emphasis
- **Match cuts / whip pans** instead of plain cuts between scenes

*Images*
- **Ken Burns effect** (slow pan/zoom on stills) so static images don't feel dead in a video
- A consistent color grade across everything so a pile of stock/AI images doesn't look disconnected
- Subtle parallax (subject separated from background) if you want something more premium than a flat zoom

*Animation*
- **Kinetic typography** for quote or list-style content
- Simple motion graphics — arrows, highlight circles, icon pop-ins — to direct attention instead of making the viewer hunt for it
- Full character/3D animation is rarely worth the time for short-form unless it's your actual differentiator

Tools: **CapCut** (fastest, built-in auto-captions/grain/light-leak presets — best for short-form), **DaVinci Resolve** (free, best color tools, steeper curve), **Premiere Pro + After Effects** (most flexible, most overhead).

**Best post-effects to layer in**

What you pasted is standard, legitimate practice — not a hidden trick:
- **Film grain** breaks up the "too clean/digital" look and reads as more intentional/filmic
- **Light leaks** add a warm, nostalgic feel — overused in some niches now, so use sparingly
- **Motion blur** smooths fast cuts and zooms so they don't feel choppy

Beyond that list: **sound design** (whooshes on transitions, a subtle ambient bed under voiceover, a hit sound on emphasis words) does more for perceived production value than any visual filter. One caution — grain, leaks, and blur are seasoning, not the meal. Stack all three on every clip and it starts reading as "a filter pack" rather than a style, which can paradoxically make AI content feel *more* templated, not less. Match the effect to the tone instead of running one preset stack on everything.

**What I think of this material**

Nothing here is fabricated — it's a real set of tactics. But my take splits by section:

- The pacing advice below is genuinely solid, widely-used editing logic — closer to industry consensus than a "secret."
- The post-effects tips are just craft. No notes.
- The "Fake BTS" hook is the one I'd actually flag. More below.

**Secret 1 — "Script Density" Retention Scale**

What it means: one idea per ~15 seconds of script (cramming multiple points into one breath makes the message muddy and unmemorable), plus a forced visual change — cut, pan, crop, caption swap, new image — every 7–10 seconds, independent of what the script is doing.

Why it works: attention habituates to an unchanging frame within a few seconds. A periodic visual "reset" recaptures it before a thumb starts drifting toward scroll, and this maps directly onto how short-form algorithms weight average view duration and rewatches.

Use it for: any feed-based short-form content competing for scroll attention — TikTok, Reels, Shorts, paid social.

Don't use it for: long-form interviews or podcasts, tutorials where the viewer needs to track one continuous demonstration without a cut breaking their place, slower or atmospheric pieces, or audiences (young kids, some neurodivergent viewers) who find rapid cutting overstimulating rather than engaging. Cutting every 8 seconds in a 20-minute cooking tutorial makes it harder to follow, not easier to watch.

**Secret 2 — "Fake Behind-the-Scenes" Hook**

What it means: platforms increasingly run classifiers that flag/label AI-generated video, so splicing in the "creation process" — screen recordings of typing prompts, dragging sliders — makes the final piece read as more human-made, while doubling as a curiosity hook.

Where I'd push back: the stated purpose is "bypasses filter algorithms" — that's evasion of an AI-disclosure system, not just an engagement device, and those are two different things:
- Showing your process because it's genuinely part of the value (an "AI art tutorial" channel, a "watch me build this" series) is a completely normal, honest hook. Fine.
- Doing it *specifically* so the content dodges an AI label or reads as more authentic/human than it is means the technique's job is to make the audience believe something false about what they're watching. That's a trust issue with your viewers, not just a clever edit — and it's a fragile bet besides, since platforms keep closing exactly this kind of loophole, and getting caught tends to cost more (penalties, credibility) than the extra watch-time bought you.

So: legitimate as a style choice, worth being skeptical of as a disclosure-dodge.
====================

That line is really the organizing principle for all of this — the skill is picking the right kit of color/transitions/effects for the video in front of you, not having one favorite stack you run on everything. Here's the fuller breakdown, built around when to reach for something and when to leave it alone.

## Color Grading

**Teal & orange** — push skin tones warm, shadows and backgrounds toward teal. Works because the two sit near-opposite on the color wheel, so it creates automatic subject/background separation.Use for: action, travel, product reveals, anything that needs energy. Avoid on talking-head or testimonial content — an obvious grade there can read as "trying too hard" and quietly undercut trust in what's being said.

**Bleach bypass (gritty/desaturated)** — pulled saturation, crushed blacks, higher contrast.Cold, serious, high-stakes. Use for: true crime, history/war content, "hard truths" exposés. Avoid it on anything meant to feel light or aspirational — this look reads heavy even laid under an upbeat script.

**Warm nostalgic / faded film** — warm color temperature, lifted (faded, not true) blacks, soft highlights. Use for: memory content, family stories, coming-of-age narratives. Avoid on fast informational content — the softness fights with sharp on-screen text.

**Clean commercial / high-key** — bright exposure, low contrast, neutral-to-cool whites. This is the Apple-ad look. Use for: tech reviews, SaaS demos, minimalist/productivity content. Avoid when you want the video to feel gritty or authentic — this grade reads as polished, sometimes to the point of corporate.

**Moody teal/blue night grade** — crushed shadows pushed cool, with selective color pops (neon, streetlights) left alone. Use for: thriller, mystery, cyberpunk-adjacent content. Avoid on daytime lifestyle or food content — it drains warmth exactly where you need appetite appeal.

**Vibrant/saturated pop** — boosted saturation, punchy contrast. Use for: food, fashion, dance, anything that needs to feel alive. Avoid on serious or vulnerable topics — oversaturation there reads as tone-deaf.

**Black & white** — full desaturation, used as a moment, not a default. Use for a single flashback, a gut-punch line, an emotional peak you want visually set apart. Avoid running it across a whole video — it stops meaning anything once it's the wallpaper instead of the accent.

## Transitions

- **Hard cut** — no transition, just a straight cut. This should be ~90% of your edits; invisible when timed to a beat or natural motion.
- **J-cut / L-cut** — next clip's audio starts before its picture (or the reverse). Use for dialogue, interviews, storytime narration — makes conversation feel continuous instead of chopped.
- **Whip pan** — a fast, blurred pan on the outgoing clip matched by a pan on the incoming one, hiding the cut inside the blur. Use for scene/location changes, high-energy vlogs, comedic beat changes. Avoid on slow or reflective content, where the sudden motion feels jarring.
- **Match cut** — cutting between two shots that share a shape, framing, or motion so the eye reads continuity instead of a break (a fist becomes a sunrise). Use for creative storytelling. Takes more planning than anything else on this list.
- **Zoom transition** — punch into full-frame on the outgoing clip, zoom out from full-frame on the incoming one. Use for chapter breaks in listicle or hook-driven short-form.
- **Glitch / digital transition** — RGB split or datamosh-style break. Use for tech, gaming, "something's wrong" narrative beats. Avoid on anything that needs to feel calm or trustworthy.
- **Light leak / flash transition** — a warm or white flash frame between cuts. Use for nostalgic montages, passage of time. It's overused in the aesthetic-vlog niche right now, so treat it as a rare accent, not a default.
- **Cross dissolve** — the classic fade-through. Use for passage of time or softening a heavy emotional cut. Reads as dated on fast short-form; better suited to longer-form or documentary pacing.
- **Speed-ramp transition** — slow motion ramping to full speed across the cut. Use for sports, dance, action highlights, big reveals.

## Effects, Matched to Tone

Since effects are seasoning, the "meal" is your story, pacing, and color choice — this table is about picking a pinch of the right one, not stacking every seasoning you own.

| Effect | Reach for it when... | Skip it when... |
|---|---|---|
| Film grain | Nostalgic, cinematic, true crime, emotional storytelling | Clean tech/tutorial content — grain reads as "lower quality" next to a crisp screen recording |
| Light leaks | Romantic, summery, memory-driven content | Serious/somber topics, or on every single clip regardless of subject |
| Motion blur | Fast cuts, action, dance, sports | Static interviews, or anywhere fine detail needs to be read (charts, UI) |
| Chromatic aberration | Retro/VHS aesthetic, glitch, horror | Clean corporate or product content |
| Vignette | Drawing the eye to a centered subject, dramatic emphasis | Wide educational shots where the frame's edges matter (whiteboards, diagrams) |
| Lens flare | Epic, inspirational, "reveal" moments | Anything meant to feel understated — overused, it turns cheesy fast |
| Deliberate camera shake | Raw, documentary urgency | Polished brand or product content |

The tell that you've over-applied: if you can't name your video's tone in one word without also listing three effects to "cover" it, you've built a filter pack, not a style.

## Cut Rhythm, by Situation

This is the practical, cut-to-cut version of the Script Density idea from before:

- **Hook (first 3 seconds)** — pattern interrupt, a cut every 1–2 seconds, bold on-screen text. This is the one place fast cutting is almost never "too fast."
- **Motivational/mindset** — cut every 3–5 seconds early, ease to 7–10 seconds mid-video, tighten back up at the call-to-action.
- **Storytime/narrative** — don't cut on a timer at all. Cut on new information or a reveal, and hold on emotional beats even if that breaks the "every 7–10 seconds" rule. Rhythm follows the story, not the clock.
- **Listicle/educational** — strict and metronomic: a new visual on every fact or number, captions locked tightly to speech.
- **True crime/mystery** — slow build, then a sharp tightening at the reveal or twist. The pacing change is itself the signal that something matters.
- **Comedy** — cut immediately after the punchline lands; don't let a joke breathe. Silence after a joke reads as a mistake, not a beat.
- **Tutorials/how-to** — minimal cutting, longer holds. The viewer needs to track one continuous action without losing their place.

## Putting It Together: Ready-to-Use Kits

| Content Type | Color Grade | Transitions | Effects | Cut Rhythm |
|---|---|---|---|---|
| Motivational/mindset | Warm-neutral, slightly desaturated for gravity | Mostly hard cuts, one light-leak flash at the turning point | Subtle grain, vignette to focus | 3–5s early → 7–10s mid → tight at CTA |
| Storytime/narrative | Warm nostalgic (flip cooler for the twist) | J-cuts for narration, whip pan for scene jumps | Light grain for a "memory" feel | Cut on story beats, not the clock |
| Listicle/educational | Clean/high-key, one LUT throughout | Zoom transitions between points | Minimal — motion graphics (arrows, circles) over filters | New visual every fact, tight caption sync |
| True crime/mystery | Bleach bypass or moody teal | Hard cuts, one glitch/static transition at the reveal | Heavy grain, vignette | Slow build, sharp tighten at the twist |
| Comedy/meme | Vibrant, timing matters more than grade | Hard cut on punchline, occasional zoom punch | Minimal — sound design matters more here | Fastest pacing, cut on every beat change |
| Tech/product showcase | Clean/high-key, neutral-cool | Smooth cross-dissolve or slide between angles | Minimal to none — clarity over style | Match cuts on rotation, longer holds to let the product be seen |
-------------------------
Focus on hook attention:
Two quick grounding facts before the formula, since "viral" is really just retention math in 2026: platforms measure hook strength in under 2 seconds now — TikTok wants the hook landing within roughly the first 1–1.3 seconds and Reels within about 1.5–2.1 seconds,timing is critical—viral hooks must capture attention within 1.3 seconds on TikTok and 2.1 seconds on Instagram Reels and anything that keeps 60%+ of viewers past the 3-second mark is what actually triggers the algorithm to push a video further.videos retaining at least 60 percent of viewers past the three-second mark are significantly more likely to be pushed to the For You Page, Reels feed, or Shorts shelf Also worth knowing: negative/"mistake" framing ("you're doing this wrong") consistently beats positive framing by a wide margin on hook rate,Negative framings consistently produce 1.3–1.8× higher hook rate on TikTok than positive framings of the same idea which matters a lot for feature-demo content since "old slow way vs. app way" is a natural mistake-frame.
==========================
Review and what do you think of this order logic?

Part 1 — Shot List
14 shots across ~48 seconds. The discipline here is exactly right:
- **Shots 1–9, 11–13** → AI-generated (photoreal video)
- **Shot 10** → Screen recording (Promedic1 app, live capture — smart; keeps clinical UI authentic and legally clean)
- **Shot 14** → Motion-graphics outro (After Effects/Canva — correct call; AI generators hallucinate logo lockups badly)

The **Locked Style Block** and **Master Negative Prompt** as fixed, copy-paste-unchanged strings is the single most important discipline decision in the whole system. Drift almost always starts when someone paraphrases the style block from memory.

### Part 2 — Consistency System
The 9-step workflow is solid. A few things worth emphasizing:

**Step 5 (reference propagation)** is underrated. Going back to the master reference sheet every shot is fine for the first 3–4 shots, but by Shot 9 you'll get tighter temporal continuity if you export a clean approved frame from Shot 3 and use *that* as the anchor for Shot 5 — same character, same lighting context, same moment in the scene.

**Step 7 (performance-driven close-ups)** is the right call for Shot 13 specifically. Text-prompting a line delivery almost always produces either a static face or a hallucinated mouth shape that doesn't match Arabic phonemes. Runway Act-Two or Kling Omni's video-reference mode with a real recorded performance will give you 10× more control on that closing shot.

**The meta-prompt for LLM shot-list generation** is structurally sound. The key mechanism — giving the agent an explicit escape hatch (`[NEEDS INPUT: ...]`) instead of letting it improvise — is exactly what prevents hallucinated props and invented character details from polluting your prompts.

### Part 3 — 9:16 Sharp-Core / Blurred-Fill
The safe-zone spec (200px top / 300px bottom / 60px sides on a 1080×1920 canvas) is correct for current TikTok/Instagram Reels UI. One addition worth noting: **compose your AI keyframes at 16:9 or wider with the subject centered** — you already have this in the tips, but it's worth making it a hard rule in the brief to whoever is prompting, not just a time-saver note. It also means you never have to re-generate a shot because the subject was too close to the edge.

### Part 4 — Tool Notes
Kling 3.0 Subject Library / Element Binding and Runway Gen-4 References are the right anchors. One flag: **verify the Omni-tier video reference feature is active on your account before you build the Shot 13 workflow around it** — it's gated by subscription tier and the UI sometimes shows the option greyed out even when it's technically available.

### Part 5 — Expert Tips
The two most important ones in practice:
1. **Render a textless master + captioned version separately** — this is a zero-cost decision at render time that saves significant rework when you release the English cut.
2. **Never let AI hallucinate on-screen medical numbers** — this isn't just a visual quality issue; it's a clinical accuracy and liability issue. The spec (keep abstract in generation, add as clean graphic overlay in post) is the correct professional standard.



---

## CLI-Agent Expert Editor — Image-to-Video Production Code

> **Purpose:** Self-contained, copy-paste-ready code that lets any CLI agent (AGY, Grok, Claude, Kimi, etc.) convert static images into broadcast-quality vertical videos with the effects described above. No external Python dependencies beyond ffmpeg/ffprobe on PATH.

### Hard Rules (non-negotiable)

1. **`-framerate 30`** on every image input (VFR trap)
2. **`-ar 48000`** on every audio output
3. **`apad` + `-shortest`** when muxing audio to video
4. **H.264** default codec (`libx264`) — HEVC only with `--archive`
5. **Never invent duration** — always `ffprobe` first
6. **Three-state QA:** PASS / FAIL / INCONCLUSIVE — INCONCLUSIVE blocks "done"

---

### 1. Core Shell Functions (zsh/bash — source once)

```bash
#!/usr/bin/env bash
# img2vid-toolkit.sh — source this file, then call functions directly
# ponytail: single file, no framework, ffmpeg-only

# ── Probe ────────────────────────────────────────────────────────────────
ffprobe_duration() {
  # Usage: ffprobe_duration <file>
  ffprobe -v error -show_entries format=duration \
    -of default=noprint_wrappers=1:nokey=1 "$1"
}

ffprobe_resolution() {
  # Returns WxH
  ffprobe -v error -select_streams v:0 \
    -show_entries stream=width,height \
    -of csv=s=x:p=0 "$1"
}

# ── Ken Burns (slow zoom-in on still image) ──────────────────────────────
# Usage: kb_zoom_in <image> <duration_sec> <output.mp4>
# ponytail: zoompan is the stdlib answer; no Python needed
kb_zoom_in() {
  local img="$1" dur="$2" out="$3"
  local frames=$((dur * 30))
  ffmpeg -y -framerate 30 -loop 1 -i "$img" \
    -vf "scale=8000:-1,zoompan=z='min(zoom+0.0015,1.5)':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d=${frames}:s=1080x1920:fps=30" \
    -t "$dur" -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -an "$out"
}

# ── Ken Burns (slow zoom-out from center) ────────────────────────────────
kb_zoom_out() {
  local img="$1" dur="$2" out="$3"
  local frames=$((dur * 30))
  ffmpeg -y -framerate 30 -loop 1 -i "$img" \
    -vf "scale=8000:-1,zoompan=z='if(eq(on,1),1.5,max(zoom-0.0015,1.0))':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d=${frames}:s=1080x1920:fps=30" \
    -t "$dur" -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -an "$out"
}

# ── Ken Burns (slow pan left→right) ──────────────────────────────────────
kb_pan_lr() {
  local img="$1" dur="$2" out="$3"
  local frames=$((dur * 30))
  ffmpeg -y -framerate 30 -loop 1 -i "$img" \
    -vf "scale=8000:-1,zoompan=z='1.2':x='(iw-iw/zoom)*on/${frames}':y='ih/2-(ih/zoom/2)':d=${frames}:s=1080x1920:fps=30" \
    -t "$dur" -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -an "$out"
}

# ── Ken Burns (slow pan right→left) ──────────────────────────────────────
kb_pan_rl() {
  local img="$1" dur="$2" out="$3"
  local frames=$((dur * 30))
  ffmpeg -y -framerate 30 -loop 1 -i "$img" \
    -vf "scale=8000:-1,zoompan=z='1.2':x='(iw-iw/zoom)*(1-on/${frames})':y='ih/2-(ih/zoom/2)':d=${frames}:s=1080x1920:fps=30" \
    -t "$dur" -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -an "$out"
}

# ── Zoom Punch (quick digital zoom burst, ~0.5s) ────────────────────────
# Usage: zoom_punch <input.mp4> <timestamp_sec> <output.mp4>
zoom_punch() {
  local inp="$1" ts="$2" out="$3"
  local punch_dur=0.5
  ffmpeg -y -i "$inp" \
    -vf "zoompan=z='if(between(in_time,${ts},${ts}+${punch_dur}),min(zoom+0.08,1.4),max(zoom-0.08,1.0))':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d=1:s=1080x1920:fps=30" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -c:a copy "$out"
}

# ── 9:16 Safe-Zone Fit (any image → vertical with blurred fill) ──────────
# Usage: fit_916 <image> <duration_sec> <output.mp4>
# ponytail: the blurred-fill technique from Part 3 of the doc
fit_916() {
  local img="$1" dur="$2" out="$3"
  ffmpeg -y -framerate 30 -loop 1 -i "$img" -loop 1 -i "$img" \
    -filter_complex "\
      [0:v]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,boxblur=20:5[bg];\
      [1:v]scale=1080:1920:force_original_aspect_ratio=decrease[fg];\
      [bg][fg]overlay=(W-w)/2:(H-h)/2" \
    -t "$dur" -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -an "$out"
}

# ── Color Grade Presets (via LUT-style eq/colorbalance filters) ──────────
# Usage: grade_<preset> <input.mp4> <output.mp4>
# ponytail: ffmpeg native filters, no external LUT files needed

grade_teal_orange() {
  ffmpeg -y -i "$1" \
    -vf "colorbalance=rs=0.15:gs=-0.05:bs=-0.15:rh=0.1:gh=0.0:bh=-0.1:rm=0.05:gm=-0.02:bm=-0.08,eq=saturation=1.2:contrast=1.05" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$2"
}

grade_bleach_bypass() {
  ffmpeg -y -i "$1" \
    -vf "eq=saturation=0.5:contrast=1.4:brightness=-0.05,curves=m='0/0 0.25/0.1 0.5/0.45 0.75/0.8 1/1'" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$2"
}

grade_warm_nostalgic() {
  ffmpeg -y -i "$1" \
    -vf "colortemperature=temperature=5500,eq=saturation=0.85:contrast=0.95,curves=m='0/0.05 0.5/0.5 1/0.95'" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$2"
}

grade_clean_highkey() {
  ffmpeg -y -i "$1" \
    -vf "eq=brightness=0.06:contrast=0.9:saturation=0.9,colortemperature=temperature=7000" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$2"
}

grade_moody_teal() {
  ffmpeg -y -i "$1" \
    -vf "colorbalance=rs=-0.1:gs=-0.05:bs=0.15:rm=-0.08:gm=0.0:bm=0.1,eq=contrast=1.3:brightness=-0.08,curves=m='0/0 0.15/0.02 0.5/0.4 1/0.9'" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$2"
}

grade_vibrant_pop() {
  ffmpeg -y -i "$1" \
    -vf "eq=saturation=1.5:contrast=1.15,unsharp=5:5:0.5:5:5:0" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$2"
}

# ── Post-Effects ─────────────────────────────────────────────────────────

# Film grain overlay (ffmpeg native noise filter)
add_grain() {
  local inp="$1" out="$2" strength="${3:-12}"
  ffmpeg -y -i "$inp" \
    -vf "noise=alls=${strength}:allf=t+u" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$out"
}

# Vignette
add_vignette() {
  local inp="$1" out="$2"
  ffmpeg -y -i "$inp" \
    -vf "vignette=PI/4" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$out"
}

# Light leak (warm flash overlay via color curves + blend)
add_light_leak() {
  local inp="$1" out="$2" intensity="${3:-0.15}"
  ffmpeg -y -i "$inp" \
    -vf "curves=r='0/0 0.5/${intensity} 1/1':g='0/0 0.5/$(echo "${intensity}*0.6" | bc) 1/1',eq=brightness=0.02" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$out"
}

# Motion blur (minterpolate-based)
add_motion_blur() {
  local inp="$1" out="$2"
  ffmpeg -y -i "$inp" \
    -vf "tblend=average,framestep=1" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy "$out"
}

# ── Transitions (xfade between two clips) ────────────────────────────────
# Usage: xfade_transition <clip1.mp4> <clip2.mp4> <output.mp4> <type> [duration]
# Types: fade, wipeleft, wiperight, wipeup, wipedown, slideleft, slideright,
#        circlecrop, dissolve, pixelize, diagtl, diagtr, hlslice, hrslice
xfade_transition() {
  local c1="$1" c2="$2" out="$3" type="${4:-fade}" xdur="${5:-0.5}"
  local c1_dur
  c1_dur=$(ffprobe_duration "$c1")
  local offset
  offset=$(echo "$c1_dur - $xdur" | bc)
  ffmpeg -y -i "$c1" -i "$c2" \
    -filter_complex "\
      [0:v][1:v]xfade=transition=${type}:duration=${xdur}:offset=${offset}[v];\
      [0:a][1:a]acrossfade=d=${xdur}[a]" \
    -map "[v]" -map "[a]" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -c:a aac -b:a 256k -ar 48000 "$out"
}

# ── Caption Burn-In (ASS/SRT subtitle overlay) ──────────────────────────
# Usage: burn_captions <input.mp4> <subtitles.srt> <output.mp4>
burn_captions() {
  local inp="$1" subs="$2" out="$3"
  # ponytail: force_style sets word-by-word animated caption aesthetic
  ffmpeg -y -i "$inp" \
    -vf "subtitles='${subs}':force_style='FontName=Arial,FontSize=22,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,Outline=2,Shadow=1,Alignment=2,MarginV=120'" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -c:a copy "$out"
}

# ── Audio Mux (image-derived video + voiceover) ─────────────────────────
# Usage: mux_audio <video.mp4> <audio.mp3> <output.mp4>
# Follows the apad + shortest rule for A/V sync
mux_audio() {
  local vid="$1" aud="$2" out="$3"
  ffmpeg -y -i "$vid" -i "$aud" \
    -c:v copy \
    -c:a aac -b:a 256k -ar 48000 \
    -af "apad" -shortest \
    "$out"
}

# ── Speed Ramp (slow→fast at a timestamp) ────────────────────────────────
# Usage: speed_ramp <input.mp4> <slow_start> <slow_end> <slow_factor> <output.mp4>
# slow_factor: 0.5 = half speed, 2.0 = double speed
speed_ramp() {
  local inp="$1" ss="$2" se="$3" factor="$4" out="$5"
  local pts_factor
  pts_factor=$(echo "1/$factor" | bc -l)
  ffmpeg -y -i "$inp" \
    -vf "setpts='if(between(T,${ss},${se}),PTS*${pts_factor},PTS)'" \
    -af "atempo=${factor}" \
    -c:v libx264 -crf 18 -pix_fmt yuv420p \
    -c:a aac -ar 48000 "$out"
}
```

---

### 2. Full Pipeline Assembler (Python — zero deps beyond stdlib + ffmpeg)

```python
#!/usr/bin/env python3
"""
img2vid_pipeline.py — Expert editor pipeline for CLI agents.
Converts a folder of images + optional audio into a polished vertical video.

Usage:
    python3 img2vid_pipeline.py --images ./shots/ --audio voice.mp3 --output final.mp4
    python3 img2vid_pipeline.py --images ./shots/ --output silent.mp4 --duration 5
    python3 img2vid_pipeline.py --images ./shots/ --audio voice.mp3 --output final.mp4 \
        --grade teal_orange --effect grain --transition fade --captions subs.srt

ponytail: one file, stdlib only, ffmpeg subprocess calls. No opencv, no pillow.
Ceiling: sequential ffmpeg calls (not piped); upgrade path is ffmpeg filter_complex graph.
"""

import argparse
import json
import os
import subprocess
import sys
import tempfile
from pathlib import Path
from typing import List, Optional, Tuple

# ── Constants ────────────────────────────────────────────────────────────

CANVAS = (1080, 1920)  # 9:16 vertical
FPS = 30
CRF = 18
CODEC = "libx264"
PIX_FMT = "yuv420p"
AUDIO_RATE = 48000

KB_EFFECTS = {
    "zoom_in":  "zoompan=z='min(zoom+0.0015,1.5)':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d={frames}:s={w}x{h}:fps={fps}",
    "zoom_out": "zoompan=z='if(eq(on,1),1.5,max(zoom-0.0015,1.0))':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d={frames}:s={w}x{h}:fps={fps}",
    "pan_lr":   "zoompan=z='1.2':x='(iw-iw/zoom)*on/{frames}':y='ih/2-(ih/zoom/2)':d={frames}:s={w}x{h}:fps={fps}",
    "pan_rl":   "zoompan=z='1.2':x='(iw-iw/zoom)*(1-on/{frames})':y='ih/2-(ih/zoom/2)':d={frames}:s={w}x{h}:fps={fps}",
}

GRADES = {
    "teal_orange":    "colorbalance=rs=0.15:gs=-0.05:bs=-0.15:rh=0.1:bh=-0.1:rm=0.05:bm=-0.08,eq=saturation=1.2:contrast=1.05",
    "bleach_bypass":  "eq=saturation=0.5:contrast=1.4:brightness=-0.05,curves=m='0/0 0.25/0.1 0.5/0.45 0.75/0.8 1/1'",
    "warm_nostalgic":  "colortemperature=temperature=5500,eq=saturation=0.85:contrast=0.95,curves=m='0/0.05 0.5/0.5 1/0.95'",
    "clean_highkey":  "eq=brightness=0.06:contrast=0.9:saturation=0.9,colortemperature=temperature=7000",
    "moody_teal":     "colorbalance=rs=-0.1:bs=0.15:rm=-0.08:bm=0.1,eq=contrast=1.3:brightness=-0.08,curves=m='0/0 0.15/0.02 0.5/0.4 1/0.9'",
    "vibrant_pop":    "eq=saturation=1.5:contrast=1.15,unsharp=5:5:0.5:5:5:0",
}

EFFECTS = {
    "grain":       "noise=alls=12:allf=t+u",
    "grain_heavy":  "noise=alls=25:allf=t+u",
    "vignette":    "vignette=PI/4",
    "light_leak":  "curves=r='0/0 0.5/0.15 1/1':g='0/0 0.5/0.09 1/1',eq=brightness=0.02",
}

XFADE_TYPES = [
    "fade", "wipeleft", "wiperight", "wipeup", "wipedown",
    "slideleft", "slideright", "circlecrop", "dissolve",
    "pixelize", "diagtl", "diagtr", "hlslice", "hrslice",
]


# ── Helpers ──────────────────────────────────────────────────────────────

def run(cmd: List[str], check: bool = True) -> subprocess.CompletedProcess:
    """Run command, print on failure."""
    result = subprocess.run(cmd, capture_output=True, text=True)
    if check and result.returncode != 0:
        print(f"FAIL: {' '.join(cmd)}", file=sys.stderr)
        print(result.stderr, file=sys.stderr)
        sys.exit(1)
    return result


def probe_duration(path: Path) -> float:
    r = run(["ffprobe", "-v", "error", "-show_entries", "format=duration",
             "-of", "default=noprint_wrappers=1:nokey=1", str(path)])
    return float(r.stdout.strip())


def probe_resolution(path: Path) -> Tuple[int, int]:
    r = run(["ffprobe", "-v", "error", "-select_streams", "v:0",
             "-show_entries", "stream=width,height",
             "-of", "csv=s=x:p=0", str(path)])
    w, h = r.stdout.strip().split("x")
    return int(w), int(h)


def collect_images(folder: Path) -> List[Path]:
    """Sorted list of image files in folder."""
    exts = {".jpg", ".jpeg", ".png", ".webp", ".bmp", ".tiff"}
    imgs = sorted(p for p in folder.iterdir() if p.suffix.lower() in exts)
    if not imgs:
        print(f"FAIL: No images found in {folder}", file=sys.stderr)
        sys.exit(1)
    return imgs


# ── Stage 1: Image → Individual Video Clips ─────────────────────────────

def image_to_clip(
    img: Path,
    duration: float,
    output: Path,
    kb_effect: str = "zoom_in",
    grade: Optional[str] = None,
    effect: Optional[str] = None,
) -> Path:
    """Convert a single image to a video clip with Ken Burns + optional grade/effect."""
    frames = int(duration * FPS)
    w, h = CANVAS

    # Build video filter chain
    vf_parts = [f"scale=8000:-1"]

    # Ken Burns motion
    kb_template = KB_EFFECTS.get(kb_effect, KB_EFFECTS["zoom_in"])
    vf_parts.append(kb_template.format(frames=frames, w=w, h=h, fps=FPS))

    # Color grade
    if grade and grade in GRADES:
        vf_parts.append(GRADES[grade])

    # Post-effect
    if effect and effect in EFFECTS:
        vf_parts.append(EFFECTS[effect])

    vf = ",".join(vf_parts)

    cmd = [
        "ffmpeg", "-y",
        "-framerate", str(FPS),  # VFR trap: always explicit
        "-loop", "1", "-i", str(img),
        "-vf", vf,
        "-t", str(duration),
        "-c:v", CODEC, "-crf", str(CRF), "-pix_fmt", PIX_FMT,
        "-an",
        str(output),
    ]
    run(cmd)
    return output


# ── Stage 2: Concatenate Clips with Transitions ─────────────────────────

def concat_clips_simple(clips: List[Path], output: Path) -> Path:
    """Lossless concat (no transitions) via demuxer."""
    list_file = output.parent / "_concat_list.txt"
    with open(list_file, "w") as f:
        for c in clips:
            f.write(f"file '{c}'\n")
    run(["ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", str(list_file),
         "-c", "copy", str(output)])
    list_file.unlink(missing_ok=True)
    return output


def concat_with_xfade(
    clips: List[Path],
    output: Path,
    transition: str = "fade",
    xfade_dur: float = 0.5,
) -> Path:
    """Chain xfade transitions across N clips."""
    if len(clips) < 2:
        return concat_clips_simple(clips, output)

    if transition not in XFADE_TYPES:
        print(f"WARNING: Unknown transition '{transition}', falling back to 'fade'", file=sys.stderr)
        transition = "fade"

    # Build massive filter_complex for N clips
    inputs = []
    for i, c in enumerate(clips):
        inputs.extend(["-i", str(c)])

    # Calculate durations
    durations = [probe_duration(c) for c in clips]

    # Build xfade chain
    fc_parts = []
    current = "[0:v]"
    for i in range(1, len(clips)):
        offset = sum(durations[:i]) - xfade_dur * i
        if offset < 0:
            offset = 0
        next_label = f"[v{i}]" if i < len(clips) - 1 else "[vout]"
        fc_parts.append(
            f"{current}[{i}:v]xfade=transition={transition}:duration={xfade_dur}:offset={offset:.3f}{next_label}"
        )
        current = next_label

    cmd = ["ffmpeg", "-y"] + inputs + [
        "-filter_complex", ";".join(fc_parts),
        "-map", "[vout]",
        "-c:v", CODEC, "-crf", str(CRF), "-pix_fmt", PIX_FMT,
        "-an",
        str(output),
    ]
    run(cmd)
    return output


# ── Stage 3: Mux Audio ──────────────────────────────────────────────────

def mux_audio(video: Path, audio: Path, output: Path) -> Path:
    """Mux audio with apad + shortest to prevent A/V drift."""
    cmd = [
        "ffmpeg", "-y",
        "-i", str(video),
        "-i", str(audio),
        "-c:v", "copy",
        "-c:a", "aac", "-b:a", "256k", "-ar", str(AUDIO_RATE),
        "-af", "apad", "-shortest",
        str(output),
    ]
    run(cmd)
    return output


# ── Stage 4: Burn Captions ───────────────────────────────────────────────

def burn_captions(video: Path, srt: Path, output: Path) -> Path:
    """Burn SRT/ASS subtitles into video with styled word captions."""
    # ponytail: MarginV=120 keeps captions above the 300px bottom safe zone
    style = "FontName=Arial,FontSize=22,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,Outline=2,Shadow=1,Alignment=2,MarginV=120"
    cmd = [
        "ffmpeg", "-y",
        "-i", str(video),
        "-vf", f"subtitles='{srt}':force_style='{style}'",
        "-c:v", CODEC, "-crf", str(CRF), "-pix_fmt", PIX_FMT,
        "-c:a", "copy",
        str(output),
    ]
    run(cmd)
    return output


# ── Stage 5: QA Gate ─────────────────────────────────────────────────────

def qa_gate(video: Path) -> str:
    """Three-state QA: check file exists, has video+audio streams, duration > 0."""
    if not video.exists():
        return "FAIL: output file does not exist"

    size = video.stat().st_size
    if size < 1024:
        return f"FAIL: output too small ({size} bytes)"

    try:
        dur = probe_duration(video)
    except Exception:
        return "FAIL: cannot probe duration"

    if dur <= 0:
        return "FAIL: zero duration"

    # Check streams
    r = run(["ffprobe", "-v", "error", "-show_entries", "stream=codec_type",
             "-of", "csv=p=0", str(video)], check=False)
    streams = r.stdout.strip().split("\n") if r.stdout.strip() else []

    has_video = "video" in streams
    has_audio = "audio" in streams

    if not has_video:
        return "FAIL: no video stream"

    # Verify resolution
    w, h = probe_resolution(video)
    if w != CANVAS[0] or h != CANVAS[1]:
        return f"INCONCLUSIVE: resolution {w}x{h}, expected {CANVAS[0]}x{CANVAS[1]}"

    # Frame rate check
    r2 = run(["ffprobe", "-v", "error", "-select_streams", "v:0",
              "-show_entries", "stream=r_frame_rate",
              "-of", "default=noprint_wrappers=1:nokey=1", str(video)], check=False)
    if r2.stdout.strip():
        num, den = r2.stdout.strip().split("/")
        actual_fps = int(num) / int(den) if int(den) > 0 else 0
        if abs(actual_fps - FPS) > 1:
            return f"INCONCLUSIVE: fps={actual_fps:.1f}, expected {FPS}"

    status = "PASS" if has_audio else "PASS (silent — no audio input)"
    info = {
        "status": status,
        "duration": f"{dur:.2f}s",
        "resolution": f"{w}x{h}",
        "file_size": f"{size / 1024 / 1024:.1f}MB",
        "streams": streams,
    }
    print(f"QA: {json.dumps(info, indent=2)}")
    return status


# ── Main Pipeline ────────────────────────────────────────────────────────

def main():
    parser = argparse.ArgumentParser(
        description="Expert image-to-video pipeline for CLI agents",
        formatter_class=argparse.RawDescriptionHelpFormatter,
        epilog="""
Examples:
  # Basic: images → video with Ken Burns, 4s per image
  python3 img2vid_pipeline.py --images ./shots/ --output reel.mp4

  # With audio, grade, effect, and transitions
  python3 img2vid_pipeline.py --images ./shots/ --audio voice.mp3 \\
    --grade teal_orange --effect grain --transition fade --output final.mp4

  # With captions burned in
  python3 img2vid_pipeline.py --images ./shots/ --audio voice.mp3 \\
    --captions subs.srt --output captioned.mp4

  # Custom duration per image, zoom out effect
  python3 img2vid_pipeline.py --images ./shots/ --duration 6 \\
    --kb-effect zoom_out --output slow.mp4

Available grades:  teal_orange, bleach_bypass, warm_nostalgic, clean_highkey, moody_teal, vibrant_pop
Available effects: grain, grain_heavy, vignette, light_leak
Available KB:      zoom_in, zoom_out, pan_lr, pan_rl
Available transitions: fade, wipeleft, wiperight, dissolve, circlecrop, pixelize, etc.
        """,
    )
    parser.add_argument("--images", required=True, help="Directory of source images")
    parser.add_argument("--audio", help="Audio/voiceover file (mp3/wav/aac)")
    parser.add_argument("--output", required=True, help="Output video path")
    parser.add_argument("--duration", type=float, default=4.0, help="Seconds per image (default: 4)")
    parser.add_argument("--kb-effect", default="zoom_in", choices=list(KB_EFFECTS.keys()), help="Ken Burns motion type")
    parser.add_argument("--kb-alternate", action="store_true", help="Alternate KB effects per image (zoom_in/zoom_out/pan_lr/pan_rl)")
    parser.add_argument("--grade", choices=list(GRADES.keys()), help="Color grade preset")
    parser.add_argument("--effect", choices=list(EFFECTS.keys()), help="Post-effect preset")
    parser.add_argument("--transition", default="none", help=f"Transition type: none, or one of {XFADE_TYPES}")
    parser.add_argument("--xfade-dur", type=float, default=0.5, help="Transition duration in seconds")
    parser.add_argument("--captions", help="SRT/ASS subtitle file to burn in")
    parser.add_argument("--archive", action="store_true", help="Use HEVC instead of H.264")
    args = parser.parse_args()

    # Override codec for archive mode
    global CODEC
    if args.archive:
        CODEC = "libx265"

    images_dir = Path(args.images)
    output = Path(args.output)
    images = collect_images(images_dir)

    # Determine per-image duration
    per_img_dur = args.duration
    if args.audio:
        audio_path = Path(args.audio)
        total_audio_dur = probe_duration(audio_path)
        per_img_dur = total_audio_dur / len(images)
        print(f"Audio: {total_audio_dur:.1f}s / {len(images)} images = {per_img_dur:.1f}s each")

    # KB effect rotation
    kb_cycle = list(KB_EFFECTS.keys())

    # Stage 1: Convert each image to a clip
    tmpdir = Path(tempfile.mkdtemp(prefix="img2vid_"))
    clips: List[Path] = []
    for i, img in enumerate(images):
        kb = kb_cycle[i % len(kb_cycle)] if args.kb_alternate else args.kb_effect
        clip_path = tmpdir / f"clip_{i:03d}.mp4"
        print(f"[{i+1}/{len(images)}] {img.name} → {kb}, {per_img_dur:.1f}s")
        image_to_clip(img, per_img_dur, clip_path, kb_effect=kb,
                       grade=args.grade, effect=args.effect)
        clips.append(clip_path)

    # Stage 2: Concatenate with optional transitions
    assembled = tmpdir / "assembled.mp4"
    if args.transition and args.transition != "none":
        print(f"Assembling {len(clips)} clips with '{args.transition}' transitions...")
        concat_with_xfade(clips, assembled, args.transition, args.xfade_dur)
    else:
        print(f"Assembling {len(clips)} clips (hard cut)...")
        concat_clips_simple(clips, assembled)

    # Stage 3: Mux audio (if provided)
    if args.audio:
        print("Muxing audio (apad + shortest)...")
        muxed = tmpdir / "muxed.mp4"
        mux_audio(assembled, Path(args.audio), muxed)
    else:
        muxed = assembled

    # Stage 4: Burn captions (if provided)
    if args.captions:
        print("Burning captions...")
        captioned = tmpdir / "captioned.mp4"
        burn_captions(muxed, Path(args.captions), captioned)
        final_src = captioned
    else:
        final_src = muxed

    # Copy to final output
    output.parent.mkdir(parents=True, exist_ok=True)
    run(["cp", str(final_src), str(output)])

    # Stage 5: QA Gate
    print("\n─── QA Gate ───")
    result = qa_gate(output)
    if result.startswith("FAIL") or result.startswith("INCONCLUSIVE"):
        print(f"❌ {result}", file=sys.stderr)
        sys.exit(1)
    else:
        print(f"✅ {result}")
        print(f"\nOutput: {output}")
        print(f"Temp files: {tmpdir} (safe to delete)")


if __name__ == "__main__":
    main()
```

---

### 3. Quick-Reference: Agent Decision Matrix

When a CLI agent receives images and needs to produce a video, use this lookup:

| Content Type | KB Effect | Grade | Effect | Transition | Cut Rhythm |
|---|---|---|---|---|---|
| Motivational/mindset | `zoom_in` | `warm_nostalgic` | `grain` | `fade` | 3–5s per image |
| Storytime/narrative | `--kb-alternate` | `warm_nostalgic` | `grain` | `dissolve` | Match story beats |
| Listicle/educational | `pan_lr` | `clean_highkey` | none | `wipeleft` | 3–4s per fact |
| True crime/mystery | `zoom_in` (slow) | `bleach_bypass` or `moody_teal` | `grain_heavy` + `vignette` | `fade` | 5–8s build, 2s at twist |
| Comedy/meme | `zoom_out` | `vibrant_pop` | none | hard cut | 2–3s per beat |
| Tech/product | `pan_lr` | `clean_highkey` | none | `slideleft` | 4–5s per angle |
| Medical/pharma | `zoom_in` | `clean_highkey` | `vignette` | `fade` | 4s per image |

**Agent command examples:**

```bash
# Motivational reel from 8 images + voiceover
python3 img2vid_pipeline.py --images ./shots/ --audio voice.mp3 \
  --kb-effect zoom_in --grade warm_nostalgic --effect grain \
  --transition fade --captions subs.srt --output motivational_reel.mp4

# Listicle with clean look
python3 img2vid_pipeline.py --images ./facts/ --audio narration.mp3 \
  --kb-effect pan_lr --grade clean_highkey --transition wipeleft \
  --duration 3.5 --output listicle.mp4

# Silent slideshow (no audio) with alternating Ken Burns
python3 img2vid_pipeline.py --images ./photos/ --duration 5 \
  --kb-alternate --grade vibrant_pop --transition dissolve \
  --output slideshow.mp4

# Quick single-image video (shell function)
source img2vid-toolkit.sh
kb_zoom_in hero.png 8 hero_clip.mp4
grade_teal_orange hero_clip.mp4 graded.mp4
add_grain graded.mp4 final.mp4 15
mux_audio final.mp4 voice.mp3 output.mp4
```

---

### 4. Self-Check (assert-based verification)

```bash
#!/usr/bin/env bash
# img2vid_selfcheck.sh — run after any pipeline execution
# ponytail: smallest thing that fails if the logic breaks

set -euo pipefail

OUTPUT="${1:?Usage: img2vid_selfcheck.sh <output.mp4>}"

echo "── Self-Check: $OUTPUT ──"

# 1. File exists and is non-trivial
[ -f "$OUTPUT" ] || { echo "FAIL: file missing"; exit 1; }
SIZE=$(stat -f%z "$OUTPUT" 2>/dev/null || stat -c%s "$OUTPUT")
[ "$SIZE" -gt 10240 ] || { echo "FAIL: file too small (${SIZE}B)"; exit 1; }

# 2. Has video stream
VCODEC=$(ffprobe -v error -select_streams v:0 -show_entries stream=codec_name -of csv=p=0 "$OUTPUT")
[ -n "$VCODEC" ] || { echo "FAIL: no video stream"; exit 1; }

# 3. Resolution is 1080x1920
RES=$(ffprobe -v error -select_streams v:0 -show_entries stream=width,height -of csv=s=x:p=0 "$OUTPUT")
[ "$RES" = "1080x1920" ] || { echo "INCONCLUSIVE: resolution=$RES, expected 1080x1920"; exit 1; }

# 4. FPS ≈ 30
FPS_RAW=$(ffprobe -v error -select_streams v:0 -show_entries stream=r_frame_rate -of default=noprint_wrappers=1:nokey=1 "$OUTPUT")
FPS_NUM=${FPS_RAW%%/*}
FPS_DEN=${FPS_RAW##*/}
FPS_ACTUAL=$((FPS_NUM / FPS_DEN))
[ "$FPS_ACTUAL" -ge 29 ] && [ "$FPS_ACTUAL" -le 31 ] || { echo "INCONCLUSIVE: fps=${FPS_ACTUAL}"; exit 1; }

# 5. Duration > 0
DUR=$(ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 "$OUTPUT")
DUR_INT=${DUR%%.*}
[ "$DUR_INT" -gt 0 ] || { echo "FAIL: zero duration"; exit 1; }

echo "PASS — ${VCODEC}, ${RES}, ${FPS_ACTUAL}fps, ${DUR}s, $(echo "$SIZE/1024/1024" | bc)MB"
```

---

### Integration Notes for Agents

1. **Save the Python script** as `img2vid_pipeline.py` in the project directory
2. **Save the shell toolkit** as `img2vid-toolkit.sh` and `source` it
3. **Save the self-check** as `img2vid_selfcheck.sh` and run after every render
4. **No pip install needed** — stdlib Python + ffmpeg on PATH is all you need
5. The pipeline auto-divides audio duration across images — just pass `--audio` and it handles timing
6. Always run `img2vid_selfcheck.sh <output>` before claiming "done" — three-state QA is mandatory
7. For face blur before conversion, use the existing `blur_faces()` from `video_generator.py` or ffmpeg's `drawbox`/external cascade



