LLMVideoPipeline/implementation_plan.md
2026-08-16 09:39:36 -06:00

11 KiB

YouTube Factory — Complete 24/7 Faceless Video Pipeline Implementation Plan

Build a comprehensive, production-ready, distributed faceless video generation pipeline utilizing 3 self-hosted PCs (AMDLLM, NvidiaLLM, IntelLLM) capable of running autonomously 24/7.

Architecture Overview

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                                 ORCHESTRATOR & 24/7 ENGINE                             │
│  • Topic Modes: Single (--topic), Batch (--batch), Live Trending (--trending)          │
│  • 24/7 Daemon Worker (--daemon) with Queue Manager & Fault-Tolerant Checkpoints       │
│  • YouTube Uploader Readiness (YouTube Data API v3 OAuth2 + Metadata + Captions)       │
└────────────────────────────────────────────────────────────────────────────────────────┘
                                     │
         ┌───────────────────────────┼───────────────────────────┐
         ▼                           ▼                           ▼
┌──────────────────┐       ┌──────────────────┐       ┌──────────────────┐
│ AMDLLM (port 8002)│       │NvidiaLLM (port 8188│       │IntelLLM (port 8003│
│ 10.4.0.181       │       │ 10.4.0.180       │       │ 10.4.0.182       │
├──────────────────┤       ├──────────────────┤       ├──────────────────┤
│ • llama.cpp LLM  │       │ • ComfyUI        │       │ • Kokoro TTS     │
│ • Qwen3.5-7B     │       │ • MiniMax H3 T2V │       │   (EN & ES)      │
│ • Scriptwriting  │       │ • z-image-turbo  │       │ • faster-whisper │
│ • Translation    │       │ • SDXL Checkpoint│       │   subtitles      │
│ • SEO & Chapters │       │ • A/B Thumbnails │       │ • FFmpeg Engine  │
│ • Trending Ideas │       │ • Stock Media    │       │ • Audio Ducking  │
└──────────────────┘       └──────────────────┘       └──────────────────┘

User Review Required

Important

Stock Media Sources: The pipeline will fetch free stock videos & images from Wikimedia Commons / Open Archives by default without any API keys required. It will also support optional Pexels and Pixabay API keys if configured in factory_config.yaml.

Important

Kokoro Server Update on IntelLLM: We will enhance kokoro_server.py on IntelLLM (:8003) to dynamically support both English (lang_code='a') and Spanish (lang_code='e') pipelines for native pronunciation across both language tracks.


Open Questions

Note

None blocking. The pipeline will support all 3 execution modes out-of-the-box:

  1. User-supplied single topic (--topic "...")
  2. User-supplied batch list (--batch "..." "...")
  3. Autonomous 24/7 trending mode (--daemon --auto-trending)

Proposed Changes

Configuration & Core

[MODIFY] factory_config.yaml

  • Add configurations for:
    • visual_strategy: hybrid (MiniMax H3 + Stock Video + AI Image + Stock Image), all_ai, or fast_stock
    • stock_media: Pexels/Pixabay API keys (optional), Wikimedia Commons enabled
    • music: procedural cinematic soundtrack generator settings, mood/genre selection, ducking volume (-20dB)
    • trending: Google Trends RSS / YouTube viral niche exploration settings
    • daemon: 24/7 interval, queue auto-refill, maximum daily video quota
    • youtube_api: client secrets path, privacy status (private, unlisted, public), auto-upload flag

Pipeline Modules (pipeline/)

[NEW] pipeline/script_gen.py

  • Generates structured documentary scripts via AMDLLM llama.cpp (Qwen3.5-7B).
  • Generates 3+ high-CTR English titles & 3+ Spanish titles for A/B testing.
  • Generates SEO descriptions with hashtags and structured chapter headings.
  • Translates narration into natural Latin American Spanish.
  • Extracts visual generation prompts and stock media search keywords for each scene.

[NEW] pipeline/tts_gen.py

  • Generates voiceover audio for English (af_heart, am_adam) and Spanish (ef_dora, em_alex) via IntelLLM Kokoro :8003.
  • Tracks per-segment exact duration and timestamps to compute precise chapter markers (00:00 - Intro, etc.).

[NEW] pipeline/visual_gen.py

  • ComfyUI MiniMax H3 Video Generator: Sends text-to-video requests to NvidiaLLM.
  • ComfyUI z-image-turbo / SDXL Image Generator: Fast high-resolution image rendering on NvidiaLLM.
  • Stock Media Downloader: Queries Wikimedia Commons and Pexels/Pixabay for relevant HD video clips and images.
  • Ken Burns Motion FX: Automatically animates still images with slow zoom/pan motion in FFmpeg.
  • Intelligent Fallback Hierarchy: MiniMax H3 Video → Stock Video → AI Image → Stock Image.

[NEW] pipeline/music_gen.py

  • Generates mood-tailored cinematic background music matching the exact target duration.
  • Supports procedural multi-layer ambient/cinematic audio synthesis (pads, harmonic chord progressions, sub-bass, atmospheric textures).
  • Generates clean looped background music ready for dynamic audio ducking.

[NEW] pipeline/sub_gen.py

  • Runs Faster-Whisper on IntelLLM to generate word-accurate .srt subtitles in English and Spanish.

[NEW] pipeline/thumb_gen.py

  • Generates 3 distinct thumbnail concepts (A, B, C) for A/B testing on NvidiaLLM.
  • For each concept, renders:
    1. English title overlay (thumb_A_en.png)
    2. Spanish title overlay (thumb_A_es.png)
    3. Clean/raw background without text (thumb_A_clean.png)
  • High-contrast typography with bold outline and dark backing box for maximum CTR.

[NEW] pipeline/assembler.py

  • Orchestrates multi-track video assembly on IntelLLM FFmpeg:
    • Video track (normalized 1080p/720p 25fps clips + Ken Burns transitions)
    • Voiceover track
    • Background music track with active audio ducking (sidechaincompress)
    • Subtitle burning and embedding
    • Embedded chapter metadata
  • Real-time Google Trends RSS query & YouTube trending feed parser.
  • AMDLLM LLM viral documentary topic ideation engine for automated 24/7 topic discovery across niches (Science, History, Tech, Space, Mysteries).

[NEW] pipeline/queue_manager.py

  • Persistent SQLite/JSON queue manager for 24/7 autonomous scheduling.
  • Tracks job statuses: QUEUED, PROCESSING, COMPLETED, FAILED.
  • Resumable checkpointing: saves intermediate artifacts so interrupted jobs resume without repeating expensive AI generation steps.

[NEW] pipeline/youtube_uploader.py

  • YouTube Data API v3 upload module:
    • Google OAuth2 authentication flow (client_secrets.json / token storage).
    • Video upload with Title, Description (including chapters), Tags, Category.
    • Custom thumbnail assignment.
    • Multi-language caption (.srt) uploading.
    • Privacy status setting (private, unlisted, public).

Main Orchestrator & CLI

[MODIFY] youtube_factory.py

  • Refactor into the unified master CLI connecting all pipeline components.
  • CLI arguments:
    • --topic "<topic>": Single video pipeline
    • --batch "<t1>" "<t2>" / --batch-file <file>: Batch video processing
    • --trending [--niche <niche>] [--count <N>]: Auto-pull trending topics
    • --daemon: 24/7 continuous autonomous worker
    • --queue-add, --queue-list, --queue-retry
    • --health: Check all 3 PCs & services
    • --upload <job_id>: YouTube API uploader
    • --dry-run: Test script & metadata generation without heavy rendering

[MODIFY] README.md

  • Update documentation with 24/7 daemon usage, stock media options, trending topic commands, background music, chapter formatting, and YouTube API upload guide.

Verification Plan

Automated & Integration Tests

  1. Health Check:
    python3 youtube_factory.py --health
    
  2. Trending Topic Discovery Test:
    python3 -c "from pipeline.trending import TrendingFetcher; print(TrendingFetcher().get_trending_topics(count=3))"
    
  3. Procedural Background Music Generation Test:
    python3 -c "from pipeline.music_gen import MusicGen; print(MusicGen().generate(30, '/tmp/test_music.wav'))"
    
  4. Stock Media Downloader Test:
    python3 -c "from pipeline.visual_gen import StockMediaFetcher; print(StockMediaFetcher().search_and_download('Ancient Rome', '/tmp/test_stock.jpg'))"
    
  5. Dry Run End-to-End Pipeline Test:
    python3 youtube_factory.py --topic "The Secrets of the Mariana Trench" --dry-run
    
  6. Full Video Generation Test (Single Topic):
    python3 youtube_factory.py --topic "The Lost Library of Alexandria" --duration 60
    
    • Verify output directory contains:
      • _en.mp4 & _es.mp4 with audio ducking and subtitles
      • _thumb_A_en.png, _thumb_A_es.png, _thumb_B_en.png, etc.
      • _meta.json with chapters, SEO description, and tags
      • _upload.txt with formatted YouTube Studio copy-paste bundle
      • _bundle.json for automated YouTube API uploader