Build me a personal, self-hosted AI video creation platform called “Creator Studio”. This is NOT a kids-only website. It should be a general-purpose video creation tool that I can use for YouTube, Shorts, Reels, TikTok, educational videos, faceless videos, storytelling, documentaries, gaming content, nursery rhymes, and any other type of video. The most important requirement is: NO PAID API DEPENDENCIES by default. Design the platform around free and open-source models and local processing. API integrations can be optional, but the application must work without paid APIs. MAIN WORKFLOW: SCRIPT → VIDEO I should be able to paste any script and automatically turn it into a complete video. Input: - Paste script - Upload TXT/DOCX/PDF - Or enter a simple topic and let AI create the script Then automatically: 1. Analyze the script 2. Divide it into scenes 3. Estimate duration 4. Generate scene descriptions 5. Generate visual prompts 6. Generate voiceover 7. Generate or select visuals 8. Add background music 9. Add sound effects 10. Generate subtitles 11. Automatically synchronize everything 12. Assemble the final video 13. Export MP4 VIDEO TYPES: Provide presets: - YouTube video - YouTube Shorts - Instagram Reels - TikTok - Educational video - Story video - Faceless video - Kids animation - Documentary - Motivational video - Podcast clips - Custom Allow custom: - Aspect ratio - Resolution - FPS - Duration - Style SUPPORTED FORMATS: - 16:9 - 9:16 - 1:1 - 4:5 AI SCRIPT GENERATOR: Create a powerful script generator where I enter: Topic: Audience: Video length: Language: Tone: Style: Generate: - Hook - Full script - Narration - Scene breakdown - Dialogue - Visual instructions - CTA Allow me to edit the generated script before creating the video. AI SCENE GENERATOR: Automatically divide scripts into scenes. Each scene should contain: - Scene number - Duration - Narration - Visual description - Image prompt - Video prompt - Camera movement - Transition - Sound effect - Background music - Subtitle text Allow: - Regenerate scene - Edit scene - Delete scene - Duplicate scene - Reorder scene - Regenerate visual - Change duration TEXT TO SPEECH: Integrate free/local TTS options such as Piper TTS, Coqui TTS or other suitable open-source TTS engines. Allow: - Male voices - Female voices - Different languages - Different accents where available - Speaking speed - Pitch - Voice preview Keep the system modular so additional TTS providers can be added later. AI IMAGE GENERATION: Support local/open-source image generation where possible. Allow: - Generate image from prompt - Generate multiple variations - Select best image - Regenerate - Upload my own image - Image-to-image - Maintain character/style consistency AI VIDEO GENERATION: Provide a video-generation module designed to support locally available/open-source models when the user's computer has sufficient GPU resources. The system should support pluggable video generation backends rather than being locked to one provider. Allow: - Text-to-video - Image-to-video - Animate image - Camera movement - Scene duration - Motion strength If local video generation is unavailable because of hardware limitations, gracefully fall back to image-based video creation using: - Generated images - Ken Burns effect - Zoom - Pan - Camera movement - Transitions VIDEO EDITOR: Create a proper timeline-based video editor. Timeline tracks: 1. Video 2. Images 3. Voiceover 4. Music 5. Sound effects 6. Subtitles 7. Text 8. Overlays Features: - Cut - Split - Trim - Crop - Resize - Rotate - Speed control - Volume control - Fade in/out - Transitions - Text overlays - Subtitles - Stickers - Images - Audio - Video layers Use FFmpeg for video processing and rendering. CLIPPING TOOL: This is VERY IMPORTANT. Create an automatic clipping system where I can upload a long video and generate short clips from it. Input: - MP4/MKV/MOV/WebM - YouTube video file - Podcast - Lecture - Interview - Long-form video Automatically: 1. Transcribe the video 2. Detect important/high-engagement moments 3. Find potential clips 4. Suggest clip start/end times 5. Generate titles 6. Generate captions 7. Convert clips to 9:16 8. Add animated subtitles 9. Reframe the speaker 10. Export multiple Shorts Show suggested clips like: CLIP 01 00:14:32 → 00:15:18 “Most interesting moment” Score: 92% CLIP 02 00:28:10 → 00:29:03 “Strong hook” Score: 88% Allow me to preview, edit and export each clip. AUTO CAPTIONS: Use free/open-source speech recognition such as Whisper or faster-whisper. Features: - Automatic transcription - Word-level timestamps - Subtitle generation - SRT/VTT export - Animated captions - Highlight important words - Multiple subtitle styles Allow caption presets: - Minimal - Bold - Shorts - Karaoke - Highlight - Professional AUDIO: Provide: - Background music - Upload music - Sound effects - Audio trimming - Volume adjustment - Noise reduction - Voice enhancement - Audio ducking Use royalty-free/open-source music by default and clearly indicate the source/license when applicable. THUMBNAIL GENERATOR: Generate YouTube thumbnails from: - Video - Screenshot - Prompt Allow: - Text - Images - Background - AI-generated elements - Multiple variations - 1280×720 export AI CONTENT PACKAGE: After creating a video automatically generate: YouTube title YouTube description SEO keywords Hashtags Tags Thumbnail text Short description Social media captions PROJECT MANAGEMENT: Create: - New Project - Save Project - Duplicate Project - Rename - Delete - Export - Import project Automatically save progress. MEDIA LIBRARY: Store: - Images - Videos - Audio - Voiceovers - Music - SFX - Thumbnails - Generated clips Allow search, filtering and folders. FREE/LOCAL-FIRST ARCHITECTURE: Prioritize free and open-source tools. Use where appropriate: - FFmpeg for video processing - Whisper/faster-whisper for transcription - Piper/other open-source TTS for voice - Open-source LLMs through Ollama - Stable Diffusion/Flux-compatible local image generation where hardware allows - Open-source video-generation models where hardware allows - OpenCV for video/image processing - Python for AI/video processing - FastAPI or Flask for backend - React/Next.js for frontend The application should detect which local AI tools are installed and show their status. Create a “System Status” page showing: ✓ FFmpeg ✓ Whisper ✓ Ollama ✓ TTS ✓ Image Generation ✓ Video Generation ✓ GPU ✓ Storage If a component isn't installed, show: - What it does - Whether it is optional - Installation instructions IMPORTANT: Do NOT pretend that unlimited AI video generation is free if the required model needs expensive GPU hardware. The software itself should be free/open-source and local-first wherever possible. For users without a powerful GPU, provide a lightweight mode using: - Local LLM - Local TTS - Whisper - User-provided images/videos - FFmpeg - Motion effects - Transitions - Captions This should still allow complete videos to be created without paid APIs. DASHBOARD: Create a beautiful professional dashboard with: “Create Video” “Generate Script” “Script → Video” “Video → Clips” “AI Voice” “Generate Images” “Generate Thumbnail” “Video Editor” Show: Recent Projects Storage Processing Jobs System Status DESIGN: Use a modern dark professional interface similar to professional creative software. Dark background. Clean cards. Subtle gradients. Smooth animations. Modern typography. Clear icons. Responsive design. Do NOT make it look like a children's website. Make it feel like a personal combination of: AI script generator + CapCut + Canva + AI video generator + Whisper transcription + FFmpeg video processor but completely focused on my own personal workflow. TECHNICAL REQUIREMENTS: Frontend: React or Next.js TypeScript Tailwind CSS Backend: Python FastAPI Video processing: FFmpeg AI: Modular provider architecture Database: SQLite for local installation, with an option to switch to PostgreSQL later. Storage: Local filesystem by default. Authentication: Optional for local personal use. All API keys must be stored in environment variables and NEVER exposed in frontend code. Create clean modular architecture so I can replace any AI model later. The application should run locally with: npm install npm run dev and the backend with: pip install -r requirements.txt python server.py Also provide a README containing complete installation instructions for Windows. The final result should be a REAL WORKING APPLICATION, not merely a UI mockup. Prioritize the following features in the first working version: 1. Script → Scene Breakdown 2. Script → Voiceover 3. Script → Images 4. Images + Voiceover → Video 5. Automatic subtitles 6. FFmpeg rendering 7. Video → Automatic Clips 8. Shorts/Reels 9:16 export 9. Thumbnail generation 10. YouTube metadata generation Build the project so additional AI video models can be plugged in later without rewriting the entire application.
Sign in to leave a comment
Architecture diagrams will be automatically generated when the Project Manager creates tasks for your project.
No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No completed page designs yet.
Completed design pages will appear here when they are ready to preview.
No comments yet. Be the first!