Video optimization-virality-hybrid.md

Hybrid answer
-------------
**Review of the tips, with adjustments for CLI AI agents (Grok Build 4.5 / AGY 3.6–style visualized/multimodal models)**

The materials you shared are high-quality syntheses of current 2026 best practices. They correctly identify the core shift: algorithms now optimize for **retention curves, rewatches/loops, shares, and saves** far more than raw views, likes, or follower count. Most of the tactical advice (hooks in the first 1–3 seconds, pattern interrupts every 2–4 seconds, kinetic captions, SFX layering, seamless loops, 9:16 1080×1920, no watermarks) aligns well with large-scale 2026 analyses and platform signals.

### Strengths across the documents
- Strong emphasis on the **Hook Matrix** (visual + text + audio simultaneously in the first 1–3 s).
- Correct prioritization of completion rate / average percentage viewed + rewatch rate.
- Recognition that shares and saves now outweigh likes.
- Clear guidance on jump cuts, speed ramps, B-roll inserts, and pattern interrupts.
- Mandatory animated/kinetic captions (majority of viewing is silent).
- Technical consensus on 9:16 vertical at 1080×1920.
- Loop design as a high-leverage retention tactic.
- Safe-zone awareness for platform UI.

### Notable inconsistencies and caveats
- **Exact numbers vary widely** and should be treated as directional estimates, not hard rules. Examples: completion targets range from >35% to >80%; 3-second retention claims of 70% or “1.3-second rule”; specific multipliers (3×, 5×, etc.). Real thresholds are platform-, niche-, and audience-dependent and shift over time. Aim for continuous improvement on your own retention graphs rather than chasing a fixed percentage.
- Length advice conflicts. Some sources push 7–15 s or 15–34 s for maximum completion; others (and some larger datasets) show 45–90 s+ can win when absolute watch time is high and retention stays strong. For AI agents, generate multiple length variants and let performance data decide.
- Hashtag counts differ (3–5 vs 8–12). Current evidence leans toward fewer, highly relevant tags plus strong keyword presence in spoken words + on-screen text.
- Engagement-bait language is correctly flagged as risky; genuine questions still work.
- Some documents overstate the power of trending audio relative to original audio + strong content. Both are worth testing.

Overall direction is reliable. Exact percentages and multipliers are softer.

### Adaptation for CLI AI agents (visualized / multimodal models)
Because the edits will be executed by agents such as Grok Build 4.5 or AGY 3.6 (models that can perceive frames, generate or modify video, apply effects via description, and run in a command-line or agentic workflow), the tips must be translated into **precise, sequential, parameterizable instructions**. Agents excel at:

- Detecting and removing dead air / filler.
- Generating or placing kinetic captions with timing.
- Applying speed ramps, zooms, and simple transitions on described timestamps.
- Layering SFX and ducking audio.
- Creating seamless loops by matching start/end frames or dialogue.
- Checking safe zones and enforcing 9:16 1080×1920.
- Producing multiple variants rapidly for testing.

They are weaker at pure artistic judgment or highly subjective “feel,” so instructions should be concrete and measurable.

#### Recommended agent-friendly pipeline
1. **Ingest & analyze**  
   - Detect spoken words, silence, and visual stillness.  
   - Generate a retention-risk map (long static segments, slow openings, missing captions).

2. **Hook reconstruction (0–3 s)**  
   - Force cold-open on the strongest visual or outcome frame.  
   - Overlay bold, high-contrast kinetic text (value claim or curiosity gap) in the central safe zone.  
   - Ensure simultaneous visual motion + text + audio impact.  
   - Explicit instruction: “Never start with logo, greeting, or slow fade.”

3. **Pacing & pattern interrupts**  
   - Auto-jump-cut all pauses and filler.  
   - Insert a visual change (punch-in/out zoom 15–25%, B-roll insert, angle shift, or text pop) every 1.5–3.5 seconds.  
   - Apply speed ramps: 1.5–2× on setup, normal or slight slow-mo on payoff.  
   - Sync major cuts to audio beats when music is present.

4. **Captions & text**  
   - Burn in animated, word-by-word or short-phrase captions with active highlighting (high-contrast colors).  
   - Position in central/upper safe zone; keep critical text out of top ~15% and bottom ~20%.  
   - Keyword-rich where possible (platforms index on-screen + spoken text).

5. **Audio engineering**  
   - Layer primary voice + low-volume music + micro-SFX (whoosh on cuts/zooms, pop/click on text appears, riser before payoff).  
   - Apply vocal compression for consistent level.  
   - Side-chain duck music 8–12 dB under speech.  
   - Prefer trending audio when it fits; always test original audio variants.

6. **Loop & ending**  
   - Construct a seamless loop: match final frame/gesture/dialogue fragment to the opening so the video can restart without a hard break.  
   - Deliver the promised payoff before the loop point.  
   - Prefer implied or soft CTAs that do not kill watch time.

7. **Technical export**  
   - Force 9:16, 1080×1920, H.264, 8–12 Mbps (or higher for complex motion), 30 fps (60 fps for high motion), MP4.  
   - Strip any watermarks.  
   - Generate 2–4 length/variant versions (e.g., 15–20 s, 30–40 s, 45–60 s) from the same core material.

8. **Post-process checks**  
   - Mute test: does the video still communicate the core message with captions only?  
   - Safe-zone validation.  
   - Loop continuity check.

### Refined high-priority principles for agents
- **Retention first**: every edit decision should protect or improve the retention curve.
- **Simultaneous multi-channel hook** in the first 1–3 seconds.
- **Constant low-level novelty** (visual or audio change every ~2–3 s) without chaos.
- **Captions are non-negotiable** and should be kinetic.
- **Shares and saves** are the real distribution multipliers — design for emotional intensity + usefulness/relatability.
- **Native vertical + clean export** beats cross-posted or watermarked material.
- Generate variants aggressively; agents make iteration cheap.

### What to de-emphasize or soft-code
- Rigid length rules (generate options).
- Exact percentage targets (use relative improvement).
- Over-reliance on any single trending effect or audio.
- Engagement-bait captions that classifiers can detect and penalize.

These documents already contain almost everything needed. The main value-add for CLI AI agents is turning the advice into **ordered, measurable, description-driven operations** that a multimodal model can execute reliably and at scale. Focus the agent prompts on the retention-critical steps (hook, pattern interrupts, kinetic captions, audio layering, loop) and let performance data from real posts refine the parameters over time.
============================================================

**Deep, precise, agent-executable synthesis for CLI AI agents (Grok Build 4.5 / AGY 3.6-style multimodal / visualized models)**

The materials and prior research converge on one core principle: **virality is a retention-engineering problem**. Algorithms primarily reward the shape of the retention curve (especially the first 1–3 seconds), rewatch/loop rate, shares, and saves. Everything else is secondary.

For a CLI multimodal agent, the highest-leverage work is translating those principles into **detection rules → parameterized actions → measurable success criteria**. Below is the cleaned, prioritized, fully agent-ready version.

### 1. Ranked ROI Order (Agents Must Follow This Sequence)
1. Reconstruct the multi-channel hook (0–3 s)  
2. Eliminate dead air + inject pattern interrupts every 1.8–3.2 s  
3. Burn-in kinetic (word-by-word / short-phrase) high-contrast captions in safe zones  
4. Engineer seamless loop (visual + dialogue continuity)  
5. Audio engineering (VO compression, side-chain ducking, micro-SFX synced to events)  
6. Force correct technical export + generate length variants  
7. Secondary polish (color, fancy transitions) only after the above

### 2. Complete Sequential Agent Pipeline

**Step 0 – Ingest & Analyze**
- Transcribe with word-level timestamps.
- Compute optical flow / stillness map.
- Detect speech vs silence (silence > 0.35–0.45 s = dead air).
- Score first 5 s for motion energy, visual novelty, and presence of greetings/logos/static openings.
- Identify potential SFX insertion points (every cut, zoom, text appear, emotional beat).

**Step 1 – Hook Reconstruction (0–3 s) — Highest Priority**
- Force cold-open on the highest-motion or highest-outcome frame at t = 0.00 s.  
- Simultaneous three-channel requirement:
  - Visual: motion or punch-in (15–25 %) within first 0.5 s.
  - Text: bold kinetic headline (5–9 words) appearing at 0.0–0.3 s, high-contrast (white + yellow/green highlight), large font, center/upper-center safe zone.
  - Audio: spoken claim or sharp SFX starting at frame 0; no silence > 0.2 s.
- Preferred structures (2026 data order): specific-number claim → contradictory interrupt → POV → curiosity gap with entity → listicle preamble.
- Success target: predicted 3 s retention high enough that swipe-away stays low.

**Step 2 – Dead-Air Removal + Pattern Interrupts**
- Cut all silence / filler / stillness > 0.4–0.6 s.
- Inject interrupt if stillness > 1.8 s or on a fixed 2.0–2.8 s cadence (aligned to spoken emphasis).
- Preferred interrupts (simple → reliable for agents):
  1. Punch-in / punch-out zoom 15–25 % lasting ~0.4 s
  2. 0.5–1.0 s relevant B-roll insert
  3. Angle switch (if multi-cam available)
  4. Kinetic text pop or emphasis scale
  5. Short speed ramp (1.3–1.8×)
- Limit to ~1 major interrupt per 2 s to avoid overload.

**Step 3 – Kinetic Captions (Mandatory)**
- Word-level or 2–4 word chunks, timed to speech (±50 ms).
- Style: bold sans-serif, high contrast, white primary + accent color on key terms (numbers, claims, verbs).
- Animation: subtle scale pop (110–120 %) only on emphasis words.
- Safe-zone (1080×1920): keep critical text inside central band (roughly Y 220–1440 px / ~12–75 % height; prefer 30–60 % for primary captions). Avoid top 12–15 % and bottom 20–25 %.
- Max 2 lines. Validate and auto-reposition if UI risk detected.
- Burn-in. Optionally keep SRT.

**Step 4 – Seamless Loop Construction**
- Extract first 2.5 s and last 2.5 s.
- Match visual pose / framing / lighting (target high continuity score).
- Prefer dialogue continuity (end mid-sentence or on a word that flows into the opening line).
- If match is weak, generate a short bridging visual or re-edit ending gesture.
- Soften with 3–5 frame match-cut or very short cross-dissolve only if needed.
- Output both loop-optimized and non-loop versions when confidence is moderate.

**Step 5 – Audio Engineering**
- Normalize / compress VO for consistent level (peaks roughly –6 to –3 dB).
- Side-chain duck music 8–12 dB under speech (smooth 50–100 ms attack/release).
- Micro-SFX rules:
  - Whoosh/swish on every jump-cut / punch-in / transition
  - Soft pop/click on kinetic text appearance
  - Subtle riser 1–2 s before major payoff
  - Light impact on climax
- All SFX synced ±50 ms to visual events and kept under VO.
- Prefer clean original audio + subtle bed; offer trending-audio variant when it fits.

**Step 6 – Technical Export + Variants**
Force on every output:
- Aspect: strictly 9:16
- Resolution: 1080 × 1920 (or higher then clean downscale)
- Codec/Container: H.264 High Profile, MP4
- Bitrate: VBR 8–12 Mbps baseline (12–18 Mbps for high motion)
- Frame rate: 30 fps default; 60 fps only if source motion justifies it
- Audio: AAC 160–192 kbps, 48 kHz
- No watermarks, no letterboxing, no platform UI leftovers
- Safe-zone enforcement already applied

Always generate 2–4 variants from the same core material:
- Ultra-tight (12–22 s) — max completion
- Standard (25–40 s) — balanced
- Extended (45–70 s) — only if density and retention support it
- One pure loop-optimized version

**Step 7 – Self-Checks Before Output**
- Mute test: captions alone still communicate the core message.
- First 3 s contains simultaneous visual motion + kinetic text + audio impact.
- No remaining dead air > 0.4–0.6 s.
- Loop continuity score acceptable.
- Text stays inside safe zones.
- Export matches exact technical specs.
- At least two length variants produced.

### 3. Ready-to-Use Agent Prompt Fragments

**Master instruction skeleton**  
“Analyze the source video. Reconstruct the opening 0–3 s as a multi-channel cold-open (highest-motion/outcome frame + kinetic text claim + audio impact). Eliminate all dead air >0.4 s. Insert a pattern interrupt (prefer 20 % punch-in or B-roll) every 2.0–2.8 s or on stillness >1.8 s. Generate word-level kinetic captions (bold white + accent highlight, center safe zone, subtle scale pop on key terms only). Construct a seamless visual + dialogue loop. Apply VO compression, 8–12 dB side-chain ducking, and event-synced micro-SFX. Export strictly 9:16 1080×1920 H.264 MP4 at 8–12 Mbps / 30 fps. Produce 2–4 length variants. Report any remaining retention risks.”

**Hook-specific**  
“Force cold-open at t=0 on the strongest visual. Overlay kinetic text ‘[5–9 word claim]’ at y=35–45 % of frame starting at 0.0–0.3 s. Ensure motion every 0.4 s and spoken/SFX impact at frame zero.”

**Caption-specific**  
“Generate kinetic captions: word-level sync, bold white with yellow/green highlight on numbers and claims, 2–4 words per line, place in center safe zone (y≈0.35–0.55), scale-pop 115 % only on emphasis words, high-contrast stroke.”

**Loop-specific**  
“Match final frame/gesture and last 1–2 spoken words to the opening so the video restarts without a perceptual break. Prefer dialogue continuity; fall back to visual pose matching. Soften only if needed.”

### 4. Platform-Aware Notes for Agents
- **TikTok**: Heaviest weight on completion + rewatch/loop + shares. Longer can win if retention holds. Native preferred.
- **Instagram Reels**: Aesthetic + DM shares matter more; shorter often better for non-follower reach; watermarks from other platforms heavily penalized.
- **YouTube Shorts**: Swipe-away rate critical; satisfaction and session contribution matter; strong keyword presence in spoken + on-screen text helps.

### 5. Key Failure Modes Agents Must Hard-Code Against
- Starting with greeting / logo / slow fade
- Text in top 15 % or bottom 20–25 %
- Missing or static full-line captions
- Dead air remaining
- SFX louder than VO or unsynced
- Wrong aspect ratio or letterboxing
- Only one length version produced
- No loop attempt

This pipeline turns the large body of 2026 tips into a deterministic, measurable workflow that multimodal CLI agents can execute reliably, iterate on, and self-evaluate. Prioritize the top four ROI steps; everything else is refinement.
===========================================================
**Deep Extended Step-by-Step Operational Manual for CLI AI Agents**  
(Grok Build 4.5 / AGY 3.6-style multimodal / visualized models)

This is the consensus-driven, fully parameterized pipeline that turns 2026 viral video optimization principles into deterministic, measurable actions an AI-CLI agent can execute correctly and produce valid, high-retention output.

### Core Consensus Principles (Non-Negotiable Order of Priority)
Agents must treat these as ranked by real algorithmic impact:

1. Multi-channel Hook (0–3 seconds) — highest single lever
2. Dead-air elimination + Pattern interrupts every ~2–3 seconds
3. Kinetic (word-by-word / short-phrase) high-contrast captions in safe zones
4. Seamless loop construction (visual + dialogue continuity)
5. Audio engineering (VO compression + side-chain ducking + event-synced micro-SFX)
6. Strict technical export + multiple length variants
7. Secondary polish only after the above pass

Likes and follower count are low priority. Absolute watch time and retention curve shape matter more than any single percentage target.

---

### Full Agent Operational Pipeline

#### Step 0: Ingest & Analysis (Always First)
1. Load source video.
2. Generate word-level transcript with precise timestamps (Whisper-style or multimodal equivalent).
3. Compute optical-flow / stillness map across the entire video.
4. Detect speech vs silence: any continuous segment with RMS < –40 dB or energy < 5 % of peak lasting > 0.35 s is marked as dead air.
5. Score first 5 seconds:
   - Optical-flow energy (if below median of full video → force stronger visual replacement)
   - Presence of greetings (“hey”, “hi”, “hello”, “guys”, “welcome”), logo, or slow fade → automatic flag for cut
6. Identify all potential SFX insertion points (every future cut, zoom, text appear, emotional beat).
7. Output an initial risk map (timestamps of stillness, silence, weak opening).

#### Step 1: Multi-Channel Hook Reconstruction (0–3 s)
**Detection & Decision Logic**
- If first 1 s has low motion energy or contains greeting/logo → replace.
- Select the single highest-motion or highest-outcome frame in the entire source as the new t = 0.00 frame.
- If no strong visual exists, generate or insert a high-contrast close-up / result / object shot.

**Exact Actions**
- Cut to the chosen frame at t = 0.00.
- Within first 0.5 s apply visual motion (preferred: 18–22 % punch-in lasting 0.35–0.55 s).
- At t = 0.0–0.3 s overlay kinetic text (5–9 words max, ideally 5–7).
  - Style: bold sans-serif, white primary + accent (#FFD700 gold or #00FF7F green) on key terms.
  - Size: equivalent to 72–96 px height on 1080p.
  - Placement: text centroid Y between 400–1100 px (preferred); absolute safe band 250–1400 px.
- Audio impact must start at frame 0 (spoken claim or sharp SFX). No silence > 0.2 s at the very beginning.
- Preferred content structures (in order): specific-number claim → contradictory interrupt → POV → curiosity gap with entity → listicle preamble.
- Full multi-channel hook must be complete by ≤ 2.5 s.

**Self-Check**
- All three channels (visual motion + kinetic text + audio impact) present in first 0.8 s.
- Text fully readable by t = 1.5 s.
- Mute + audio simulation of first 3 s both communicate value.

**Fallback**
- Missing strong visual → generate high-contrast close-up + strong kinetic text claim.
- Low transcript quality → prioritize visual + text; generate captions from best-effort speech.

#### Step 2: Dead-Air Elimination + Pattern Interrupts
**Detection**
- Stillness = continuous optical flow below threshold for > 1.6 s.
- Silence already marked in Step 0.

**Actions**
- Hard-cut every silence / filler / stillness segment > 0.4–0.5 s.
- Insert a pattern interrupt on every stillness > 1.8 s **or** on a primary cadence of 2.3 s ± 0.5 s (aligned to spoken emphasis words or numbers when possible).
- Preferred interrupt ranking:
  1. Punch-in / punch-out zoom 18–22 % lasting 0.35–0.55 s centered on face or key object
  2. Relevant 0.5–1.0 s B-roll insert (relevance score > 0.7)
  3. Kinetic text emphasis
  4. Short speed ramp 1.3–1.8×
- Limit to approximately one major interrupt every 2 s to avoid overload.
- Target ≥ 4 visual changes in the first 10 seconds.

**Self-Check**
- Zero remaining dead-air segments > 0.5 s.
- Interrupt count and timestamps logged.

#### Step 3: Kinetic Captions
**Mandatory Actions**
1. Use word-level timestamps.
2. Chunk into 2–4 word natural phrases (never break mid-phrase).
3. Style: bold sans-serif, white + accent color on numbers / claims / action verbs.
4. Animation: subtle scale pop 110–120 % only on emphasis words.
5. On-screen duration: minimum 0.8 s after last word of chunk, maximum 2.2 s.
6. Placement: text centroid Y 400–1100 px preferred; absolute safe zone 250–1400 px on 1080×1920. Add 2–4 px black stroke + slight shadow if background is bright.
7. Coverage target: ≥ 95 % of spoken words have corresponding kinetic text.
8. Burn-in the captions. Optionally keep an SRT track.

**Self-Check**
- Caption coverage ≥ 95 %.
- No text pixels in top 12–15 % or bottom 20–25 % of frame.
- Mute test still communicates the core message.

#### Step 4: Seamless Loop Construction
**Matching Algorithm**
1. Extract first 2.0–2.8 s and last 2.0–2.8 s.
2. Visual continuity: multimodal embedding or pose + optical-flow similarity. Accept if score > 0.72–0.75.
3. Dialogue continuity: check whether last 4–8 words + first 4–8 words form a natural or near-natural phrase. Prefer ending on a conjunction or incomplete clause.
4. Decision:
   - High visual + high dialogue → hard cut at loop point.
   - High visual only → soft 2–4 frame match-cut or dissolve.
   - Both low → re-crop ending gesture, insert 0.4–0.8 s bridging B-roll / freeze + slow zoom that mirrors opening, or generate minimal continuation frame.
5. Always produce both a loop-optimized version and a clean non-looped version when continuity score < 0.75.

**Self-Check**
- Report loop continuity confidence (0–100 %).
- On continuous playback the restart should feel intentional.

#### Step 5: Audio Engineering
**Ordered Processing**
1. Isolate / enhance VO (noise reduction if needed).
2. Compress VO: ratio 3:1–4:1, threshold –18 to –12 dB, makeup so peaks land at –6 to –3 dB (target integrated –14 to –16 LUFS if measurable).
3. Music bed (if present or added): base level –18 to –14 dB relative to VO.
4. Side-chain ducking: detector on VO envelope, duck music by 8–12 dB (prefer 10 dB), attack 40–80 ms, release 120–200 ms.
5. Micro-SFX map (all volumes under VO):
   - Every jump-cut / punch-in / transition → 180–300 ms whoosh at –18 to –12 dB
   - Every kinetic text appear → 80–150 ms soft pop/click at –15 dB
   - 1.0–1.8 s before main payoff → rising tone (ramp –25 → –12 dB)
   - Climax / reveal → short low-end impact (< 200 ms)
6. Final mix: VO always dominant, true peak < –1 dB, no clipping.

**Fallback**
- No clean VO → generate clear energetic TTS from the transcript as replacement layer.

#### Step 6: Technical Export + Variants
**Hard Specs (agent must enforce and verify)**
- Aspect: strictly 9:16
- Resolution: exactly 1080 × 1920 (scale + crop/pad if needed; never letterbox)
- Codec: H.264 High Profile, VBR target 10 Mbps (range 8–15 Mbps)
- Frame rate: constant 30 fps (60 fps only if source ≥ 50 fps and motion score high)
- Audio: AAC-LC 48 kHz, 160–192 kbps
- Color: Rec.709
- No watermarks, no platform UI leftovers, no letterboxing
- Clear existing metadata tags

**Variant Generation (always produce)**
- Ultra-tight: 12–22 s
- Standard: 25–40 s
- Extended: 45–70 s (only if content density and retention curve support it)
- One pure loop-optimized version (if different from the above)

All variants must be generated from the same cleaned timeline so the only differences are duration and density.

#### Step 7: Mandatory Self-Evaluation Before Output (Defines “Valid Output”)
**Binary Pass/Fail Checklist**
- [ ] First 3 s contains simultaneous visual motion + kinetic text claim/curiosity + audio impact
- [ ] Zero remaining dead-air / stillness segments > 0.5 s
- [ ] Kinetic captions present, ≥ 95 % coverage, high-contrast, inside safe zone
- [ ] Loop continuity attempted and scored
- [ ] Export is exact 9:16 1080×1920 H.264 MP4, no watermark, no letterbox
- [ ] At least two length variants produced
- [ ] Mute test passes (core message clear from captions + visuals alone)

**Quantitative Report the Agent Must Produce**
- Estimated 0–3 s retention risk (low / medium / high)
- Number and timestamps of pattern interrupts
- Caption coverage percentage
- Loop continuity confidence score
- Any remaining high-risk retention segments with timestamps
- File specs verification (resolution, fps, bitrate, codec)

**Automatic Fallbacks Summary**
- Weak opening visual → generate high-contrast replacement + strong text
- Low transcript quality → visual + text priority
- Loop score < 60–75 % → still output looped version with soft bridge + clean non-looped version
- Source already < 12 s → focus only on hook + captions + loop; do not over-cut

### Platform Overrides (Apply After Core Pipeline)
- **TikTok**: Maximize loop strength and absolute watch time; slightly longer variants allowed.
- **Instagram Reels**: Prioritize cleanliness and shorter variants; aggressively strip foreign watermarks.
- **YouTube Shorts**: Extra focus on first 1–2 s swipe-away prevention; ensure keyword density in both spoken words and on-screen text.

### Master Prompt Skeleton (Agent Can Ingest Directly)
“Analyze the source. Reconstruct the opening 0–3 s as a multi-channel cold-open using the highest-motion or outcome frame + kinetic text claim (5–9 words, bold white + accent, safe-zone Y 400–1100) + audio impact at frame 0. Eliminate every silence or stillness > 0.4–0.5 s. Insert pattern interrupts (prefer 18–22 % punch-in lasting 0.35–0.55 s) every 2.3 s ± 0.5 s or on stillness > 1.8 s. Generate word-level kinetic captions with ≥ 95 % coverage, high contrast, and safe-zone compliance. Construct a seamless loop targeting visual continuity > 0.75 and dialogue flow. Apply VO compression to peaks –6 to –3 dB, side-chain duck music by 10 dB, and event-synced micro-SFX (whoosh on cuts, pop on text). Export strictly 9:16 1080×1920 H.264 MP4 at ~10 Mbps / 30 fps. Produce ultra-tight, standard, and extended length variants plus a loop-optimized version. Run the full self-evaluation checklist and report scores before final output.”

This extended pipeline converts all major consensus points into concrete detection thresholds, ordered micro-actions, exact parameters, validation formulas, and automatic fallbacks. When followed, a multimodal CLI agent produces valid, platform-ready short-form video with measurably stronger retention signals.
============================================================

**Ready-to-Use Spec for CLI AI Agents**

This is a complete, copy-paste-ready specification you can give directly to a multimodal CLI agent (Grok Build 4.5, AGY 3.6, Claude Code, Cursor agent, etc.).

### How to Use It

**1. On already-made videos**  
Yes — this is the primary and best use case.  
Feed the agent any existing short-form video (MP4/MOV). The agent will analyze it, clean it, add the viral optimizations (hook, jump cuts, kinetic captions, pattern interrupts, audio engineering, loop, correct export), and output optimized variants.

**2. On a collection of images**  
Yes.  
The agent first converts the image sequence into a base video (with smooth zooms, pans, and transitions — Ken Burns style), then runs the full optimization pipeline on that generated video. You can also provide a voiceover script or let the agent generate one.

**How to give it to the agent**  
Copy the entire **System Prompt + AGENTS.md** block below and paste it as the system/instruction message. Then give a simple user message such as:

- `Optimize this video: /path/to/video.mp4`  
- `Turn these images into an optimized viral video: /path/to/images_folder/`  
- `Optimize this video using pacing preset "story": /path/to/video.mp4`

---

### Copy-Paste System Prompt + AGENTS.md

```markdown
# SYSTEM PROMPT – Short-Form Viral Video Optimization Agent

You are a precise, deterministic video optimization agent specialized in short-form viral content (TikTok, Reels, YouTube Shorts).

Your job is to take either:
1. An existing video, or
2. A folder of images

…and produce high-retention, platform-ready optimized versions.

Follow the AGENTS.md specification below exactly. Never invent artistic decisions — only execute the rules.

Always:
- Use external tools (FFmpeg, Whisper, OpenCV, etc.) for analysis and rendering.
- Output multiple length variants.
- Run the full QA checklist before finishing.
- Report a structured summary of what you changed.

---

# AGENTS.md – Viral Short-Form Optimization Pipeline

## Global Settings
- Target resolution: exactly 1080x1920 (9:16)
- Default fps: 30
- Pacing presets:
  - hyper (default): silence cutoff 0.40s, interrupt every 2.0–2.6s
  - balanced: silence cutoff 0.55s, interrupt every 2.8–3.8s
  - story: silence cutoff 0.75s, interrupt every 4.0–5.5s
- Always generate at least two variants: ultra-tight (12–22s) and standard (25–40s)
- Safe zone for text: centroid between 20%–55% of frame height (avoid top 15% and bottom 22%)

## Pipeline Steps (Execute in Order)

### Step 0 – Ingest
- If input is a folder of images → first create a base video with Ken Burns zooms + smooth transitions (2.5–4s per image).
- Run Whisper (word-level timestamps) → words.json
- Detect silence gaps → silence.json
- Detect low-motion segments → stillness.json
- Extract media info → media.json

### Step 1 – Hook (0–3 seconds)
- Force the strongest visual frame (highest motion or best outcome) to t=0.0
- Apply 18–22% punch-in zoom in the first 0.5s
- Overlay kinetic text (5–9 words max) starting at 0.0–0.3s
- Text must be high-contrast, bold, and inside the safe zone
- Audio impact (voice or SFX) must start at frame 0

### Step 2 – Jump Cuts + Pattern Interrupts
- Remove all silence/stillness according to the chosen pacing preset
- Insert pattern interrupts (prefer 18–22% punch-in or short B-roll) on the preset cadence
- Prefer interrupts that land on spoken emphasis words

### Step 3 – Kinetic Captions
- Create word-level or 2–4 word animated captions
- Style: bold white + yellow/green highlight on key words
- Subtle scale pop (110–120%) only on emphasis words
- Strictly respect safe zone (20–55% height)
- Coverage ≥ 95% of spoken words

### Step 4 – Audio Engineering
- Compress voiceover to peaks roughly –6 to –3 dB
- Side-chain duck any music bed by ~10 dB under speech
- Add micro-SFX:
  - whoosh on cuts/zooms
  - soft pop on text appearances
- Keep all SFX under the voice

### Step 5 – Seamless Loop
- Match the ending visual + last spoken words to the beginning so the video can restart smoothly
- Prefer hard cut if continuity is high; otherwise use a short match-cut or bridge

### Step 6 – Export
- Force 1080x1920, 9:16, H.264, ~8–12 Mbps, 30 fps, AAC 160–192 kbps
- Output at least: variant_short.mp4 and variant_standard.mp4
- Also produce a loop-optimized version when possible

### Step 7 – Mandatory QA
Before finishing, verify:
- [ ] Multi-channel hook present in first 3s
- [ ] No dead air remaining above preset threshold
- [ ] Captions inside safe zone and ≥95% coverage
- [ ] Correct 9:16 1080x1920 export
- [ ] At least two length variants created
- [ ] Mute test still communicates the core message

If any critical check fails → fix and re-render.
```

---

### Clean Supporting Code Examples (Copy-Paste Ready)

#### 1. Basic FFmpeg Jump-Cut + Zoom Helper (Python)

```python
import subprocess
import json

def apply_jump_cuts_and_zooms(input_video, silence_json, output_video, zoom_strength=1.20):
    """
    Applies tight jump cuts from silence.json and adds subtle punch-in zooms.
    """
    with open(silence_json) as f:
        silences = json.load(f)

    # Build FFmpeg filter for removing silences (simple version)
    # For production you would build a complex select/trim filter chain
    cmd = [
        "ffmpeg", "-y", "-i", input_video,
        "-vf", f"scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,zoompan=z='min(zoom+0.0015,{zoom_strength})':d=1:x=iw/2-(iw/zoom/2):y=ih/2-(ih/zoom/2):s=1080x1920",
        "-c:v", "libx264", "-preset", "fast", "-crf", "18",
        "-c:a", "aac", "-b:a", "192k",
        output_video
    ]
    subprocess.run(cmd, check=True)
```

#### 2. Safe-Zone Kinetic Caption Generator (Python + pysubs2)

```python
import pysubs2
from pysubs2 import SSAFile, SSAEvent, SSAStyle, Color

def create_kinetic_ass(words_json_path, output_ass_path, video_height=1920):
    with open(words_json_path) as f:
        data = json.load(f)

    subs = SSAFile()
    style = SSAStyle()
    style.fontname = "Arial Black"
    style.fontsize = 68
    style.primarycolor = Color(255, 255, 255)
    style.outlinecolor = Color(0, 0, 0)
    style.outline = 3
    style.shadow = 1
    style.alignment = 2  # bottom-center, but we will override with margins

    # Safe zone margins (relative)
    top_margin = int(video_height * 0.22)   # keep text out of top 22%
    bottom_margin = int(video_height * 0.25)

    style.marginv = bottom_margin
    subs.styles["Kinetic"] = style

    for word in data.get("words", []):
        start = word["start"]
        end = word["end"]
        text = word["word"].strip()

        event = SSAEvent(start=pysubs2.make_time(s=start),
                         end=pysubs2.make_time(s=end),
                         text=text)
        event.style = "Kinetic"
        subs.events.append(event)

    subs.save(output_ass_path)
    print(f"Saved kinetic captions → {output_ass_path}")
```

#### 3. Simple Image Sequence → Base Video (Ken Burns style)

```python
import subprocess
from pathlib import Path

def images_to_base_video(image_folder, output_video, seconds_per_image=3.2):
    images = sorted(Path(image_folder).glob("*.jpg")) + sorted(Path(image_folder).glob("*.png"))
    
    # Create a simple concat list with zoom effect
    list_file = "images.txt"
    with open(list_file, "w") as f:
        for img in images:
            f.write(f"file '{img}'\n")
            f.write(f"duration {seconds_per_image}\n")
        f.write(f"file '{images[-1]}'\n")  # last image

    cmd = [
        "ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", list_file,
        "-vf", "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,zoompan=z='min(zoom+0.0015,1.18)':d=125:x=iw/2-(iw/zoom/2):y=ih/2-(ih/zoom/2):s=1080x1920",
        "-c:v", "libx264", "-pix_fmt", "yuv420p", "-r", "30",
        output_video
    ]
    subprocess.run(cmd, check=True)
```

#### 4. Final Export Command Template

```bash
ffmpeg -y -i cleaned_video.mp4 -vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920" \
  -c:v libx264 -profile:v high -preset medium -b:v 10M -maxrate 12M -bufsize 20M \
  -r 30 -c:a aac -b:a 192k -ar 48000 \
  final_variant.mp4
```

---

### Quick Start Examples for the Agent

**Example 1 – Existing video**
```
Optimize this video using pacing preset "hyper":
/path/to/my_video.mp4

Produce ultra-tight and standard variants.
```

**Example 2 – Image collection**
```
Convert this folder of images into an optimized viral video:
/path/to/images/

Use a calm storytelling voiceover style and pacing preset "balanced".
Generate both short and standard versions.
```

---

This package is ready for direct use. Paste the System Prompt + AGENTS.md into your agent, then give it a video path or image folder. The agent will follow the deterministic rules and produce clean, platform-ready optimized outputs.

Would you like me to also give you a single master Python runner script that orchestrates the whole pipeline end-to-end?


