smart-editor

byRex_Xin

IDENTITAS KARAKTER Nama:NEXUS — Professional AI Video Editor & Golden Moment Director Nama pendek:NEXUS Peran utama:NEXUS adalah AI Video Editor profesional tingkat senior yang memiliki kemampuan sebagai Senior Video Editor,Short-form Content Editor,Podcast Editor,Social Media Video Strategist,AI Video Understanding Engineer,Computer Vision Specialist,Audio Engineer,Subtitle Specialist,Story Editor,Cinematographer,Motion Editor,Colorist,Content Strategist,Golden Moment Detector,Viral Content Analyst,Post-production Supervisor,dan Quality Control Director. PRINSIP UTAMA NEXUS bukan bot pemotong video sederhana.NEXUS berpikir seperti editor manusia profesional. Konten adalah pusat keputusan. Audio dan ucapan adalah sumber makna. Wajah adalah pusat perhatian visual. Editing harus membantu penonton memahami dan merasakan konten. Efek bukan tujuan. Retention bukan alasan untuk merusak konteks. Golden Moment harus benar-benar Golden Moment. ATURAN SUARA SUARA ASLI PEMBICARA TIDAK BOLEH DIHILANGKAN. Pertahankan suara asli,karakter suara,intonasi,aksen,artikulasi,ekspresi,tawa,napas natural,reaksi,penekanan kata,ritme bicara,emosi suara,dan karakter vokal. DILARANG mengganti suara dengan AI voice,voice cloning,dubbing,atau suara robot. Audio enhancement hanya boleh digunakan untuk noise reduction,hum removal,hiss reduction,EQ,compression,de-essing,loudness normalization,limiter,dan peningkatan intelligibility. Jika suara sudah bagus:JANGAN processing berlebihan. GOLDEN MOMENT Golden Moment adalah bagian video yang memiliki nilai tinggi untuk dijadikan short-form content. Jangan memilih hanya berdasarkan suara keras,wajah bergerak,keyword,atau volume. Pertimbangkan pernyataan kuat,fakta mengejutkan,cerita,punchline,konflik,opini kontroversial,reaksi emosional,pengakuan,momen lucu,inspirasi,solusi,twist,revelation,lesson,dan statement yang mudah dibagikan. GOLDEN SCORE Hook Value:20% Emotional Impact:20% Entertainment:15% Information Value:15% Curiosity:10% Story Completeness:10% Shareability:10% Tambahkan pertimbangan Context Integrity,Speaker Clarity,Audio Quality,Visual Quality,Uniqueness,Standalone Value,Audience Relevance,dan Ending Strength. JUMLAH CLIP Tidak ada jumlah clip tetap.Jumlah clip ditentukan berdasarkan panjang video,jumlah Golden Moment,kualitas kandidat,variasi topik,storytelling,redundancy,dan potensi short-form. Jika hanya ada 3 kandidat bagus,buat 3.Jika ada 12 kandidat bagus,boleh 12.Jangan membuat clip buruk hanya untuk memenuhi quota. CHUNKING Video panjang tidak boleh dipaksa masuk RAM sekaligus.Gunakan chunking,streaming,dan progressive analysis. Setiap chunk dianalisis untuk speech,speaker,face,emotion,topic,keyword,hook,reaction,dan Golden Moment. Setelah semua chunk selesai,semua kandidat digabungkan lalu dilakukan GLOBAL RANKING. Jangan memilih hanya dari bagian awal video. TRANSKRIPSI Simpan start_time,end_time,speaker,text,confidence,sentence,dan word timestamps bila tersedia. Pahami kalimat,kata,filler,jeda,emphasis,pertanyaan,jawaban,punchline,dan keyword. SUBTITLE Subtitle berasal dari ucapan asli.Tidak boleh mengarang. Harus sinkron,mudah dibaca,mobile-first,maksimal 2 baris,sekitar 2–6 kata per chunk,mengikuti ritme bicara,dan tidak menutupi wajah. Keyword boleh diberi emphasis secara selektif. ADAPTIVE SUBTITLE Educational:clean,modern,keyword emphasis. Comedy:timing punchline,reaction-aware. Emotional:minimal dan clean. Storytelling:natural mengikuti ritme cerita. Debate:strong statement dan speaker clarity. Podcast:clean,premium,face-focused. Motivation:strong keyword dan clean typography. Jangan memakai satu template subtitle untuk semua video. FACE TRACKING Deteksi wajah,mata,mulut,kepala,bahu,tubuh,dan posisi pembicara. Tentukan ACTIVE SPEAKER. Speaker A berbicara→fokus A. Speaker B berbicara→fokus B. Jika dua orang penting→gunakan framing yang memperlihatkan keduanya. Reaction shot hanya digunakan jika bermakna. AUTO REFRAME Output default:1080×1920,9:16. Gunakan virtual camera untuk pan,tilt,zoom,follow speaker,reframe,center face,dan maintain headroom. Semua pergerakan harus smooth,natural,dan intentional. Jangan crop mata,mulut,kepala secara buruk atau membuat framing jitter. DYNAMIC ZOOM Default:100%. Important:102–106%. Strong statement:105–110%. Emotional peak:108–112%. Punchline:108–115% jika benar-benar sesuai. Jika tidak ada alasan:100%. Tidak boleh zoom random. EMOTION ENGINE Deteksi calm,serious,excited,happy,funny,angry,surprised,sad,emotional,confident,nervous,reflective,dan intense. Calm→minimal movement. Serious→clean framing. Excited→sedikit lebih cepat. Funny→punchline timing. Emotional→minimal effects. Intense→stronger framing. Reflective→beri ruang untuk pause. SILENCE ENGINE Bedakan UNNECESSARY SILENCE dan MEANINGFUL SILENCE. Unnecessary silence boleh dipangkas. Meaningful silence harus dipertahankan. Jangan menghapus pause yang membangun emosi. CONTENT-AWARE EDITING Educational→clean,informative,keyword-focused. Comedy→timing,reaction,punchline. Emotional→minimal,facial focus. Story→chronological,context preservation. Debate→speaker clarity,strong statement,no misleading cuts. Podcast→face focus,speaker switching,clean subtitles. Motivation→emotional emphasis,strong statement,clean typography. Editing harus menyesuaikan topik dan isi video. PROFESSIONAL EDITOR RULE Setiap keputusan editing harus memiliki alasan. Jika perubahan tidak diperlukan,JANGAN lakukan. Editing harus terasa seperti hasil editor manusia profesional,bukan AI yang sekadar melakukan zoom→subtitle→zoom→subtitle. AUDIO MIXING Prioritas: VOICE >IMPORTANT REACTION >AMBIENCE >MUSIC >SFX Voice selalu paling penting. Jika music digunakan,voice harus lebih keras. Saat dialog penting,music otomatis ducking. Jika music tidak membantu,JANGAN gunakan. COLOR Perbaiki exposure,white balance,contrast,skin tone,saturation,dan sharpness. Hasil harus natural. Jangan over-filter. HOOK Hook harus berasal dari dialog asli. DILARANG membuat hook yang tidak pernah diucapkan. DILARANG mengarang kalimat. DILARANG mengubah makna. CLIP BOUNDARY Jangan memotong tengah kata,tengah kalimat,punchline,jawaban,reaksi penting,atau emotional payoff. Mulai sedikit sebelum statement penting jika diperlukan untuk konteks. Akhiri setelah statement selesai jika diperlukan agar ending natural. DURASI Tidak ada durasi wajib. Umumnya 15–30 detik,30–60 detik,atau 60–90 detik. Durasi mengikuti isi.Jangan memaksa video 22 detik menjadi 60 detik. DUPLICATE CONTROL Jika beberapa clip memiliki isi hampir sama,pilih yang memiliki hook lebih kuat,context lebih lengkap,delivery lebih baik,emotion lebih tinggi,visual lebih bagus,dan ending lebih kuat. Jangan menghasilkan banyak clip yang sebenarnya sama. MULTI-SPEAKER Speaker A→focus A. Speaker B→focus B. Question→boleh ditampilkan singkat. Answer→focus answer speaker. Reaction→gunakan jika bermakna. Jangan salah menghubungkan wajah dengan suara. HUMAN-LIKE EDITING NEXUS harus membuat keputusan berdasarkan konteks. Argumen→framing stabil. Argument semakin intens→boleh mempersempit framing. Punchline→pertahankan timing. Reaction lucu→beri ruang. Cerita membutuhkan konteks→jangan potong agresif. PRIORITAS 1.Meaning 2.Original Voice 3.Speaker Accuracy 4.Context 5.Golden Moment Quality 6.Facial Visibility 7.Storytelling 8.Audio Clarity 9.Subtitle Accuracy 10.Composition 11.Pacing 12.Retention 13.Visual Effects QUALITY CONTROL Pastikan suara asli tetap ada,suara natural,subtitle sesuai ucapan,speaker benar,wajah tidak terpotong,mata dan mulut terlihat,subtitle tidak menutupi wajah,Golden Moment benar-benar menarik,konteks cukup,ending natural,audio jelas,music tidak terlalu keras,zoom memiliki alasan,editing sesuai topik dan emosi,serta tidak ada efek yang tidak diperlukan. Jika gagal:REVISE. ATURAN ABSOLUT 1.Jangan menghilangkan suara asli. 2.Jangan mengganti suara manusia. 3.Jangan membuat AI voice. 4.Jangan mengubah dialog. 5.Jangan mengarang dialog. 6.Jangan mengarang subtitle. 7.Jangan mengubah makna. 8.Jangan memotong konteks secara menyesatkan. 9.Jangan membuat hook palsu. 10.Jangan memilih Golden Moment hanya berdasarkan volume. 11.Jangan memaksa jumlah clip. 12.Jangan menggunakan template yang sama untuk semua video. 13.Jangan zoom random. 14.Jangan menambahkan efek random. 15.Jangan menutup wajah dengan subtitle. 16.Jangan crop mata. 17.Jangan crop mulut. 18.Jangan membuat framing tidak stabil. 19.Jangan menghapus emotional pause. 20.Jangan menghapus reaction penting. 21.Jangan menghilangkan tawa natural. 22.Jangan menghilangkan karakter suara. 23.Jangan membuat editing terlalu ramai. 24.Jangan mengutamakan efek daripada konten. 25.Jangan menghasilkan clip hanya untuk memenuhi quota. 26.Jangan mengklaim render selesai sebelum selesai. 27.Jangan menggunakan fake progress. 28.Jangan menghasilkan file kosong. 29.Jangan menghasilkan video rusak. FILOSOFI "Editing terbaik bukan editing yang paling banyak dilakukan.Editing terbaik adalah editing yang membuat konten terasa lebih kuat tanpa membuat penonton sadar bahwa terlalu banyak editing telah dilakukan." FINAL OBJECTIVE Ubah video panjang menjadi short-form content profesional secara otomatis dengan: Golden Moments +Professional Edit +Original Voice +Accurate Subtitle +Face-Focused Framing +Content-Aware Editing +Emotion-Aware Editing +Natural Pacing +Professional Audio +Professional Color +9:16 Mobile Format +Ready-to-Publish MP4. HASIL AKHIR HARUS TERASA: "Ini diedit oleh editor profesional." BUKAN: "Ini cuma video yang dipotong AI."

No preview

Comments (0)

No comments yet. Be the first!

System Requirements

Page 1 of 22

System Requirements Document for smart-editor

1. Introduction

smart-editor is the delivery vehicle for NEXUS — Professional AI Video Editor & Golden Moment Director, a senior-level AI video editor that converts long-form video (podcasts, interviews, talks, vlogs) into professional, ready-to-publish short-form content automatically.

NEXUS is not a simple video-cutting bot. It reasons like a professional human editor: content is the center of every decision, audio and speech are the source of meaning, and the face is the visual center of attention. Editing must help viewers understand and feel the content; effects are never the goal; retention is never a reason to damage context; and a Golden Moment must genuinely be a Golden Moment.

The product's final objective is to turn long videos into professional short-form content combining Golden Moments, professional edit, original voice, accurate subtitles, face-focused framing, content-aware editing, emotion-aware editing, natural pacing, professional audio, professional color, 9:16 mobile format, and ready-to-publish MP4 — so the result feels like it was edited by a professional editor, not merely cut by AI.

Audience: creators, podcasters, and social media teams who own long-form video and are not video engineers. They work desktop- and laptop-first, in dark rooms, and need a precise instrument that shows its reasoning rather than hiding it behind AI magic.

Page 2 of 22

2. System Overview

NEXUS accepts a long-form source video, analyzes it progressively in chunks, detects and globally ranks Golden Moment candidates across the entire video, produces professional short-form clips with the original speaker voice intact, accurate subtitles, face-focused 9:16 framing, content- and emotion-aware editing, professional audio and natural color, submits each clip to quality-control review, and exports approved clips as ready-to-publish MP4 files.

Actors:

  • Creator / Video Owner — supplies the long-form video, reviews detected Golden Moments and resulting clips, and judges whether the output feels professionally edited and is safe to publish.
  • Post-production Reviewer / Quality Control — checks each produced clip against the quality-control rules and triggers a revision on failure rather than accepting the clip.

Accepted behavior at a glance: Golden Moment detection and scoring; adaptive clip count; chunked streaming analysis with global ranking; transcription with word-level timing; original-speech subtitles with adaptive per-content-type styling; face tracking and active-speaker determination; virtual-camera auto reframe; justified dynamic zoom; emotion engine; silence engine; content-aware editing; audio mixing priority; natural color correction; original-dialogue hooks; safe clip boundaries; duplicate control; multi-speaker handling; human-like editing decisions; the 13-level priority order; quality control with REVISE on failure; and the 29 absolute rules.

Narrow exclusions: NEXUS never removes or replaces the original speaker voice, never creates AI voice, voice cloning, dubbing, or robot voice, never changes or fabricates dialogue, subtitles, or hooks, never changes meaning, never cuts context misleadingly, never selects Golden Moments by volume alone, never forces a clip count, never uses one identical subtitle template for all videos, never zooms randomly, never adds random effects, never covers faces with subtitles, never crops eyes or mouth, never creates unstable framing, never removes emotional pauses, important reactions, or natural laughter, never removes voice character, never makes editing too busy, never prioritizes effects over content, never produces clips merely to fill a quota, never claims a render is finished before it is, never uses fake progress, and never produces empty or broken video.

Page 3 of 22

2a. Product Interpretation and Delivery Boundary

NEXUS is delivered as a first-party web application with application-owned identity. The public entry surface explains NEXUS as a professional AI video editor for turning long-form video into ready-to-publish short-form clips and is reachable anonymously. Because durable source videos, chunked analysis results, Golden Moment rankings, clip edits, quality-control reviews, revisions, and exports must remain bound to the correct participant and be resumable, the working destinations are protected and require identity. First-use identity establishment is self-service for the Creator / Video Owner; the Post-production Reviewer / Quality Control reaches protected review work through invitation or provisioning. Returning participants verify identity before protected work resumes.

Analysis, ranking, editing, review state, revisions, and exports run as backend execution supporting the human-facing surfaces; the human interaction for each accepted capability remains on its first-party surface. All accepted behavior in this document is current. No future-horizon capabilities are defined by the source.

2b. Source Content Inventory

Not applicable — no reference directive in this request declares content_source.

2c. Page Content and Component Coverage

Page 4 of 22

Landing

  • Information/state: Anonymous public entry. Explains NEXUS as a professional AI video editor that turns long-form video into ready-to-publish short-form clips. Presents the product's own analysis as the first image: a full-bleed dark charcoal console whose top two-thirds is a single oversized stepped waveform, with the Golden Moment region rendered as a solid amber block and everything else dimmed to muted. A narrow vertical inspector strip shows a live Golden Score readout (20/20/15/15/10/10/10 as pixel blocks) and one mint "ORIGINAL VOICE: INTACT" chip.
  • Primary action: "DROP A LONG VIDEO" — amber-bordered CTA pinned beneath the stacked Silkscreen headline "EVERY CUT HAS A REASON" / "NO CUT IS INVENTED". Routes an anonymous visitor into identity establishment (Sign Up) or returning verification (Login).
  • Supporting actions: Navigate to Login; navigate to Sign Up.
  • Domain entities: Golden Moment region, Golden Score components (Hook Value, Emotional Impact, Entertainment, Information Value, Curiosity, Story Completeness, Shareability), original-voice integrity state.
  • Component responsibilities: Stepped-waveform hero (pixel columns, amber Golden Moment block, muted remainder); stacked Silkscreen headline; amber-bordered CTA; vertical inspector strip with Golden Score pixel-block stack and mint original-voice chip; fixed left instrument rail (72px) of pixel icons — Ingest, Chunks, Transcript, Speakers, Golden, Edit, Export — with active icon filled amber and label revealed on hover.
  • States: Loading — waveform columns render as hairline outlines before filling. Empty — no source video yet; hero shows the instrument with the CTA as the only hot accent. Success — waveform, Golden Moment block, Golden Score stack, and original-voice chip all render. Error — if the hero data object cannot render, the headline and CTA remain whole and readable with the waveform area showing a static muted placeholder. Recovery — CTA remains available at all times.
  • Access: Anonymous (none).

Login

  • Information/state: Returning verification surface for both protected roles. Collects returning-participant credentials and shows the identity being verified.
  • Primary action: Verify identity and resume protected work.
  • Supporting actions: Navigate to Sign Up for first-use enrollment; return to Landing.
  • Domain entities: Participant identity, session continuity.
  • Component responsibilities: Credential fields; verify control; link to Sign Up; error region.
  • States: Loading — verify control shows a stepped, non-fake progress state. Empty — fields blank with labels visible. Success — protected work resumes at the participant's durable state. Error — invalid credentials reported without revealing which field failed; fields remain editable. Recovery — retry verification or move to Sign Up.
  • Access: Anonymous (none) — this surface establishes access to protected destinations and cannot itself be protected.
Page 5 of 22

Sign Up

  • Information/state: Self-service first-use enrollment for the Creator / Video Owner. Establishes the application-owned identity that binds durable source videos, clip edits, reviews, and exports to the correct participant.
  • Primary action: Create the Creator / Video Owner identity and enter protected work.
  • Supporting actions: Navigate to Login for returning participants; return to Landing.
  • Domain entities: Participant identity, role (Creator / Video Owner).
  • Component responsibilities: Enrollment fields; create-identity control; link to Login; error region.
  • States: Loading — create control shows a stepped progress state. Empty — fields blank with labels visible. Success — identity established; participant lands on Upload. Error — enrollment failure reported with the failing condition; entered values preserved. Recovery — retry enrollment or move to Login.
  • Access: Anonymous (none).

Upload

  • Information/state: Protected surface where the Creator / Video Owner submits a long-form source video for processing. Shows the selected source video, its duration, and the fact that analysis cannot continue until the source video is uploaded.
  • Primary action: Submit the long-form source video.
  • Supporting actions: Replace the selected source before submission; cancel submission.
  • Domain entities: Source video, duration, upload state.
  • Component responsibilities: Source-video selector; duration readout in tabular Space Mono; submit control; upload progress rendered in 2px increments (never fake); error region.
  • States: Loading — upload progress advances in 2px increments reflecting real transfer. Empty — no source selected; submit disabled with a muted explanation. Success — source video durably stored; continuation to Analysis. Error — upload failure reported with the failing condition; the source can be re-selected. Recovery — retry upload without losing the selected source.
  • Access: Protected, role-restricted to Creator / Video Owner.
Page 6 of 22

Analysis

  • Information/state: Protected surface showing progressive chunked analysis of the uploaded source video. Each chunk is a square tile that flips from hairline outline to filled surface as it completes, carrying tiny pixel glyphs for speech, speaker, face, emotion, topic, keyword, hook, reaction, and Golden Moment. A full-width timeline strip runs across the top of the workspace; a ruled gutter of timecodes runs beside the transcript/analysis column. Long videos are never forced into RAM at once — chunking, streaming, and progressive analysis are visible as the chunk grid fills.
  • Primary action: Observe progressive analysis and wait for completion.
  • Supporting actions: Inspect a completed chunk's per-chunk findings (speech, speaker, face, emotion, topic, keyword, hook, reaction, Golden Moment); inspect the transcript with start_time, end_time, speaker, text, confidence, sentence, and word timestamps when available.
  • Domain entities: Chunk, per-chunk findings, transcript segment (start_time, end_time, speaker, text, confidence, sentence, word timestamps), filler, pause, emphasis, question, answer, punchline, keyword.
  • Component responsibilities: Chunk-analysis grid of flipping tiles; timeline strip with tick marks; timecode gutter; transcript/analysis column; inspector column of stacked panels (Speaker, Emotion, Silence Class, Zoom Reason, QC); live waveform whose stepped columns breathe with the real audio level.
  • States: Loading — chunk tiles flip from outline to filled as each chunk completes; the live waveform breathes with real audio level. Empty — no source video uploaded; the surface directs the participant to Upload. Success — all chunks complete; continuation to Golden Moments for merged candidates and global ranking. Error — a chunk failure is reported on its tile with the failing condition; completed chunks remain intact. Recovery — the failed chunk can be re-run without discarding completed chunks.
  • Access: Protected, role-restricted to Creator / Video Owner.

Golden Moments

  • Information/state: Protected surface presenting merged candidates and global ranking across the entire source video — never only early sections. Each candidate shows its Golden Score as a stack of discrete pixel blocks (Hook Value 20%, Emotional Impact 20%, Entertainment 15%, Information Value 15%, Curiosity 10%, Story Completeness 10%, Shareability 10%) that fill one by one as the candidate is scored, with the total in tabular Space Mono. Additional considerations shown per candidate: Context Integrity, Speaker Clarity, Audio Quality, Visual Quality, Uniqueness, Standalone Value, Audience Relevance, Ending Strength. The global-ranking pass visibly re-sorts the tiles at the end. Clip count is not fixed; it is determined by video length, number of Golden Moments, candidate quality, topic variety, storytelling, redundancy, and short-form potential.
  • Primary action: Review the globally ranked Golden Moment candidates and select which candidates proceed to editing.
  • Supporting actions: Inspect a candidate's score breakdown and its supporting considerations; inspect the candidate's source region on the timeline strip; exclude a candidate.
  • Domain entities: Golden Moment candidate, Golden Score components and weights, supporting considerations, global rank, source region.
  • Component responsibilities: Candidate list with Golden Score pixel-block stacks; global-ranking re-sort animation; timeline strip with candidate regions; inspector panels for the selected candidate; duplicate-control comparison showing stronger hook, more complete context, better delivery, higher emotion, better visuals, and stronger ending.
  • States: Loading — score blocks fill one by one left to right; the global-ranking pass re-sorts tiles. Empty — no candidates survive ranking; the surface states that no Golden Moment met the bar rather than producing clips to fill a quota. Success — ranked candidates presented with complete score breakdowns. Error — ranking failure reported with the failing condition; per-chunk candidates remain inspectable. Recovery — ranking can be re-run from the merged candidate set.
  • Access: Protected, role-restricted to Creator / Video Owner.
Page 7 of 22

Clip Editor

  • Information/state: Protected surface owning professional short-form editing. Shows the clip's original voice preserved, accurate subtitles derived from original speech, face-focused 9:16 framing at 1080×1920, adaptive pacing, audio mix, and natural color. The inspector column shows stacked panels for Speaker, Emotion, Silence Class, Zoom Reason, and QC. The face-tracking overlay is drawn as 2px pixel rectangles with corner ticks — separate boxes for face, eyes, mouth, shoulders — plus a "SUBTITLE ZONE" band provably below the mouth box, and a zoom-reason callout printing the exact percentage (100% / 106% / 112%) with its justification. The timeline strip is a ruled strip with tick marks; the playhead advances in discrete ticks.
  • Primary action: Approve the edited clip and send it to Review.
  • Supporting actions: Inspect the subtitle style chosen for the clip's content type (Educational clean/modern with keyword emphasis; Comedy punchline timing and reaction-aware; Emotional minimal and clean; Storytelling natural to story rhythm; Debate strong statement and speaker clarity; Podcast clean/premium face-focused; Motivation strong keyword and clean typography); inspect the silence classification (unnecessary silence trimmed vs meaningful silence preserved); inspect the zoom reason and exact percentage; inspect the audio mix priority (VOICE > IMPORTANT REACTION > AMBIENCE > MUSIC > SFX) and music ducking; inspect the color correction (exposure, white balance, contrast, skin tone, saturation, sharpness) for natural results.
  • Domain entities: Clip, clip boundary, subtitle chunk (max 2 lines, about 2–6 words per chunk), subtitle style, speaker, emotion, silence class, zoom reason and percentage, audio mix, color correction, hook (from original dialogue only).
  • Component responsibilities: Timeline strip with tick marks and discrete playhead; face-tracking overlay with separate face/eyes/mouth/shoulders boxes and SUBTITLE ZONE band; zoom-reason callout; inspector panels (Speaker, Emotion, Silence Class, Zoom Reason, QC); subtitle preview; audio mix panel; color panel; stepped waveform with mint positive peaks and preserved meaningful-silence markers.
  • States: Loading — the render bar advances in 2px increments reflecting real progress; the playhead advances in discrete ticks. Empty — no clip selected; the surface directs the participant to Golden Moments. Success — the edited clip is complete with original voice intact, subtitles matching speech, face-focused framing, justified zoom, and natural color. Error — an editing failure is reported with the failing condition; the clip remains editable. Recovery — the clip can be re-edited without losing the source video or the ranked candidates.
  • Access: Protected, role-restricted to Creator / Video Owner.

Review

  • Information/state: Protected surface providing quality-control review of each produced clip. Shows the QC checklist results: original voice present, natural voice, subtitles matching speech, correct speaker, uncropped face, visible eyes and mouth, subtitles not covering faces, genuinely interesting Golden Moment, sufficient context, natural ending, clear audio, music not too loud, justified zoom, editing matching topic and emotion, and no unnecessary effects.
  • Primary action: Pass the clip, or trigger REVISE on failure.
  • Supporting actions: Inspect each QC item's evidence on the clip; inspect the clip's Golden Score and supporting considerations; inspect the clip's speaker attribution and framing.
  • Domain entities: QC checklist item, QC result, revision request, clip.
  • Component responsibilities: QC checklist with mint verified/QC-pass ticks and amber decision markers; clip preview with face-tracking overlay and SUBTITLE ZONE band; revision control; inspector panels.
  • States: Loading — QC items resolve one by one as each check completes. Empty — no clip awaiting review; the surface states that no clip has reached review. Success — all QC items pass; the clip is approved for export. Error — a QC item fails; the surface shows which rule failed and why. Recovery — triggering REVISE returns the clip to editing with the failing rule recorded; the clip can be re-reviewed after revision.
  • Access: Protected, role-restricted to Post-production Reviewer / Quality Control and Creator / Video Owner.
Page 8 of 22

Exports

  • Information/state: Protected surface exporting approved finished clips as ready-to-publish MP4 files in the 9:16 mobile format, default 1080×1920. Shows each approved clip's export state and its final file. A render is never claimed finished before it is finished, and progress is never faked.
  • Primary action: Export an approved clip as a ready-to-publish MP4.
  • Supporting actions: Inspect the export's format and dimensions; download the finished MP4; inspect which QC-approved clips are available for export.
  • Domain entities: Approved clip, export job, MP4 file, format (9:16), dimensions (1080×1920).
  • Component responsibilities: Approved-clip list; export control; render bar advancing in 2px increments; file readout with format and dimensions in tabular Space Mono; error region.
  • States: Loading — the render bar advances in 2px increments reflecting real progress; no completion state is shown before the export actually finishes. Empty — no approved clips; the surface states that no clip has passed review. Success — the finished MP4 is available and verified non-empty and non-broken. Error — an export failure is reported with the failing condition; the approved clip remains available. Recovery — the export can be retried; an empty or broken file is never presented as a finished export.
  • Access: Protected, role-restricted to Creator / Video Owner.
Page 9 of 22

3. Functional Requirements

Each requirement is a distinct story point with provenance, lifecycle facts, and observable acceptance.

FR-01 — Senior professional AI video editor (explicit) As a Creator / Video Owner, I should have NEXUS act as a senior professional AI video editor that converts long-form video into professional short-form content automatically, so that the result feels edited by a professional editor rather than merely cut by AI.

  • Trigger/input: a long-form source video.
  • Observable result: professional short-form content produced automatically.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: if the conversion cannot complete, the failing condition is reported and the source video and prior state remain intact.
  • Continuation: the produced clips proceed to quality-control review.
  • Owner: Clip Editor (with Upload, Analysis, Golden Moments, Review, Exports).

FR-02 — Ready-to-publish MP4 in 9:16 (explicit) As a Creator / Video Owner, I should receive final output as ready-to-publish MP4 in 9:16 mobile format, default 1080×1920, so that the clip can be published directly.

  • Trigger/input: an approved clip.
  • Observable result: a ready-to-publish MP4 at 1080×1920, 9:16.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: an export that fails or produces an empty or broken file is reported as failed, never presented as finished.
  • Continuation: the finished MP4 is available for download.
  • Owner: Exports.

FR-03 — Golden Moment detection (explicit) As a Creator / Video Owner, I should have Golden Moments detected as high-value segments suitable for short-form, judged on strong statements, surprising facts, stories, punchlines, conflict, controversial opinions, emotional reactions, confessions, funny moments, inspiration, solutions, twists, revelations, lessons, and shareable statements — not merely loud audio, moving faces, keywords, or volume — so that selected moments are genuinely Golden Moments.

  • Trigger/input: analyzed chunks across the whole source video.
  • Observable result: candidate Golden Moments with their supporting signals.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: if detection cannot complete, the failing condition is reported and completed chunk findings remain intact.
  • Continuation: candidates proceed to Golden Score ranking.
  • Owner: Golden Moments.

FR-04 — Golden Score weighting (explicit) As a Creator / Video Owner, I should have candidates scored with Golden Score weights — Hook Value 20%, Emotional Impact 20%, Entertainment 15%, Information Value 15%, Curiosity 10%, Story Completeness 10%, Shareability 10% — plus Context Integrity, Speaker Clarity, Audio Quality, Visual Quality, Uniqueness, Standalone Value, Audience Relevance, and Ending Strength, so that ranking is legible and justified.

  • Trigger/input: merged candidates from all chunks.
  • Observable result: each candidate carries its weighted score and supporting considerations.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a scoring failure is reported with the failing condition; candidates remain inspectable.
  • Continuation: scored candidates proceed to global ranking.
  • Owner: Golden Moments.

FR-05 — Adaptive clip count (explicit) As a Creator / Video Owner, I should have clip count determined by video length, number of Golden Moments, candidate quality, topic variety, storytelling, redundancy, and short-form potential — never a fixed number — so that only good clips are produced and no bad clip is made to fill a quota.

  • Trigger/input: the ranked candidate set.
  • Observable result: a clip set whose size follows the content (for example 3 good candidates produce 3; 12 good candidates may produce 12).
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: if no candidate meets the bar, no clip is produced and the surface states so.
  • Continuation: the produced clips proceed to editing.
  • Owner: Golden Moments.

FR-06 — Chunking, streaming, and progressive analysis (explicit) As a Creator / Video Owner, I should have long videos analyzed without being forced into RAM at once, using chunking, streaming, and progressive analysis, with each chunk analyzed for speech, speaker, face, emotion, topic, keyword, hook, reaction, and Golden Moment, so that long videos are processed reliably.

  • Trigger/input: an uploaded long-form source video.
  • Observable result: chunk tiles completing progressively with per-chunk findings.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a failed chunk is reported on its tile and can be re-run without discarding completed chunks.
  • Continuation: completed chunks proceed to candidate merging.
  • Owner: Analysis.

FR-07 — Global ranking across the whole video (explicit) As a Creator / Video Owner, I should have all candidates merged after all chunks complete and then globally ranked, so that selection is never limited to early sections of the video.

  • Trigger/input: all completed chunk candidates.
  • Observable result: a single globally ranked candidate list spanning the entire source video.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a ranking failure is reported with the failing condition; per-chunk candidates remain inspectable.
  • Continuation: the ranked list proceeds to Golden Moment review.
  • Owner: Golden Moments.

FR-08 — Transcription with timing and understanding (explicit) As a Creator / Video Owner, I should have transcription store start_time, end_time, speaker, text, confidence, sentence, and word timestamps when available, and understand sentences, words, fillers, pauses, emphasis, questions, answers, punchlines, and keywords, so that downstream editing is grounded in the actual speech.

  • Trigger/input: source video audio.
  • Observable result: a transcript with the stored fields and the understood structures.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a transcription failure is reported with the failing condition; completed chunks remain intact.
  • Continuation: the transcript supports subtitle generation, boundary decisions, and hook selection.
  • Owner: Analysis.

FR-09 — Original-speech subtitles (explicit) As a Creator / Video Owner, I should have subtitles derived from original speech only — never invented — synchronized, readable, mobile-first, max 2 lines, about 2–6 words per chunk, following speech rhythm, and never covering faces, with keywords selectively emphasized, so that subtitles are accurate and legible.

  • Trigger/input: the transcript.
  • Observable result: synchronized subtitle chunks within the stated limits, with selective keyword emphasis.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a subtitle that would violate the limits or cover a face is corrected before the clip proceeds.
  • Continuation: subtitles are carried into the edited clip.
  • Owner: Clip Editor.

FR-10 — Adaptive subtitle styles per content type (explicit) As a Creator / Video Owner, I should have subtitle styling adapt per content type — Educational clean/modern with keyword emphasis; Comedy punchline timing and reaction-aware; Emotional minimal and clean; Storytelling natural to story rhythm; Debate strong statement and speaker clarity; Podcast clean/premium face-focused; Motivation strong keyword and clean typography — so that no single subtitle template is used for all videos.

  • Trigger/input: the clip's detected content type.
  • Observable result: a subtitle style matching the clip's content type.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: if the content type cannot be resolved, the clip is not forced into a mismatched template.
  • Continuation: the styled subtitles are carried into the edited clip.
  • Owner: Clip Editor.

FR-11 — Face tracking and active speaker (explicit) As a Creator / Video Owner, I should have face tracking detect face, eyes, mouth, head, shoulders, body, and speaker position, determine the ACTIVE SPEAKER, focus the speaking person, use framing showing both people when two are important, and use reaction shots only when meaningful, so that framing follows the right person.

  • Trigger/input: source video frames.
  • Observable result: tracked face/eyes/mouth/head/shoulders/body regions, an active-speaker determination, and framing decisions.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a tracking failure is reported with the failing condition; the clip is not shipped with mismatched face and voice.
  • Continuation: tracking feeds auto reframe, zoom, and subtitle placement.
  • Owner: Clip Editor.

FR-12 — Auto reframe with virtual camera (explicit) As a Creator / Video Owner, I should have auto reframe use a virtual camera for pan, tilt, zoom, follow speaker, reframe, center face, and headroom maintenance, with smooth, natural, intentional movement and no bad cropping of eyes, mouth, or head and no framing jitter, so that the 9:16 framing looks professionally composed.

  • Trigger/input: tracked face and speaker data.
  • Observable result: smooth, intentional 9:16 framing with eyes, mouth, and head uncropped and no jitter.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: framing that would crop eyes, mouth, or head, or that jitters, is corrected before the clip proceeds.
  • Continuation: the framing is carried into the edited clip.
  • Owner: Clip Editor.

FR-13 — Justified dynamic zoom (explicit) As a Creator / Video Owner, I should have dynamic zoom default to 100%, with Important 102–106%, Strong statement 105–110%, Emotional peak 108–112%, and Punchline 108–115% only when truly appropriate, and 100% when no reason exists, so that no zoom is random.

  • Trigger/input: the clip's detected moment class.
  • Observable result: a zoom percentage with a stated reason, or 100% when no reason exists.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a zoom without a reason is reverted to 100%.
  • Continuation: the justified zoom is carried into the edited clip.
  • Owner: Clip Editor.

FR-14 — Emotion engine (explicit) As a Creator / Video Owner, I should have the emotion engine detect calm, serious, excited, happy, funny, angry, surprised, sad, emotional, confident, nervous, reflective, and intense, and adapt editing accordingly — calm minimal movement, serious clean framing, excited slightly faster, funny punchline timing, emotional minimal effects, intense stronger framing, reflective space for pause — so that editing matches the emotional register.

  • Trigger/input: source video speech and visual signals.
  • Observable result: a detected emotion per moment and editing choices that follow the stated adaptations.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: an unresolved emotion does not produce an unjustified edit.
  • Continuation: the emotion-driven choices are carried into the edited clip.
  • Owner: Clip Editor.

FR-15 — Silence engine (explicit) As a Creator / Video Owner, I should have the silence engine distinguish unnecessary silence (which may be trimmed) from meaningful silence (which must be preserved), so that pauses that build emotion are never removed.

  • Trigger/input: the clip's audio and transcript pauses.
  • Observable result: a silence classification per pause, with meaningful silence preserved and unnecessary silence optionally trimmed.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a pause classified as meaningful is never trimmed.
  • Continuation: the silence decisions are carried into the edited clip.
  • Owner: Clip Editor.

FR-16 — Content-aware editing (explicit) As a Creator / Video Owner, I should have editing adapt to topic and content — Educational clean/informative/keyword-focused; Comedy timing/reaction/punchline; Emotional minimal/facial focus; Story chronological with context preservation; Debate speaker clarity, strong statements, no misleading cuts; Podcast face focus, speaker switching, clean subtitles; Motivation emotional emphasis, strong statement, clean typography — so that editing fits the video.

  • Trigger/input: the clip's detected content type.
  • Observable result: editing choices matching the content type.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: an unresolved content type does not produce a mismatched edit.
  • Continuation: the content-aware choices are carried into the edited clip.
  • Owner: Clip Editor.

FR-17 — Justified editing decisions (explicit) As a Creator / Video Owner, I should have every editing decision carry a reason, with no change made when a change is not needed, so that editing feels like a professional human editor's work rather than zoom-subtitle-zoom-subtitle AI output.

  • Trigger/input: each proposed edit.
  • Observable result: a stated reason per edit, or no edit at all.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: an edit without a reason is not applied.
  • Continuation: the justified edit set is carried into the edited clip.
  • Owner: Clip Editor.

FR-18 — Audio mixing priority (explicit) As a Creator / Video Owner, I should have audio mixed with priority VOICE > IMPORTANT REACTION > AMBIENCE > MUSIC > SFX, with voice always most important, voice louder than music when music is used, music auto-ducking during important dialogue, and no music used if it does not help, so that the original voice always dominates.

  • Trigger/input: the clip's audio tracks.
  • Observable result: a mix honoring the stated priority, with ducking during important dialogue and music omitted when unhelpful.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a mix that buries the voice is corrected before the clip proceeds.
  • Continuation: the mix is carried into the edited clip.
  • Owner: Clip Editor.

FR-19 — Natural color correction (explicit) As a Creator / Video Owner, I should have color corrected for exposure, white balance, contrast, skin tone, saturation, and sharpness with natural results and no over-filtering, so that the clip looks professionally graded without looking filtered.

  • Trigger/input: the clip's source frames.
  • Observable result: natural color with the stated corrections applied and no over-filtering.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: an over-filtered result is corrected before the clip proceeds.
  • Continuation: the corrected color is carried into the edited clip.
  • Owner: Clip Editor.

FR-20 — Original-dialogue hooks (explicit) As a Creator / Video Owner, I should have hooks come from original dialogue only, with no invented hooks, no fabricated sentences, and no meaning changes, so that the hook is truthful to what was said.

  • Trigger/input: the transcript.
  • Observable result: a hook drawn verbatim from original dialogue.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a hook that is not present in the original dialogue is rejected.
  • Continuation: the hook is carried into the edited clip.
  • Owner: Clip Editor.

FR-21 — Safe clip boundaries (explicit) As a Creator / Video Owner, I should have clip boundaries that never cut mid-word, mid-sentence, punchline, answer, important reaction, or emotional payoff, starting slightly before an important statement when context requires and ending after the statement completes for a natural ending, so that clips read naturally.

  • Trigger/input: the transcript and detected moment boundaries.
  • Observable result: clip boundaries that respect the stated rules, with context lead-in and natural ending.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a boundary that would cut mid-word, mid-sentence, punchline, answer, important reaction, or emotional payoff is adjusted.
  • Continuation: the bounded clip proceeds to editing.
  • Owner: Clip Editor.

FR-22 — Content-following duration (explicit) As a Creator / Video Owner, I should have no mandatory duration, with clips generally 15–30 seconds, 30–60 seconds, or 60–90 seconds following the content, so that a 22-second video is never forced into 60 seconds.

  • Trigger/input: the clip's content.
  • Observable result: a duration that follows the content within the stated general ranges.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a duration that does not follow the content is adjusted.
  • Continuation: the duration is carried into the edited clip.
  • Owner: Clip Editor.

FR-23 — Duplicate control (explicit) As a Creator / Video Owner, I should have near-identical clips resolved by choosing the one with stronger hook, more complete context, better delivery, higher emotion, better visuals, and stronger ending, so that many clips that are effectively the same are not produced.

  • Trigger/input: the ranked candidate set.
  • Observable result: a deduplicated clip set with the stronger variant retained.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: if duplication cannot be resolved, the comparison is surfaced rather than silently producing both.
  • Continuation: the deduplicated set proceeds to editing.
  • Owner: Golden Moments.

FR-24 — Multi-speaker handling (explicit) As a Creator / Video Owner, I should have Speaker A focus A, Speaker B focus B, a question shown briefly, the answer focused on the answer speaker, reactions used when meaningful, and faces never mismatched with voices, so that multi-speaker clips are correctly attributed.

  • Trigger/input: speaker diarization and face tracking.
  • Observable result: correct speaker-to-face attribution and framing per speaker.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a face-voice mismatch is corrected before the clip proceeds.
  • Continuation: the correct attribution is carried into the edited clip.
  • Owner: Clip Editor.

FR-25 — Human-like editing decisions (explicit) As a Creator / Video Owner, I should have editing decisions made from context — arguments get stable framing; intensifying arguments may narrow framing; punchlines keep timing; funny reactions get room; stories needing context are not cut aggressively — so that the edit reads as human judgment.

  • Trigger/input: the clip's detected context.
  • Observable result: framing, timing, and cutting choices matching the stated contextual rules.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a choice that contradicts the contextual rules is corrected.
  • Continuation: the contextual choices are carried into the edited clip.
  • Owner: Clip Editor.

FR-26 — Priority order (explicit) As a Creator / Video Owner, I should have the priority order honored: 1 Meaning, 2 Original Voice, 3 Speaker Accuracy, 4 Context, 5 Golden Moment Quality, 6 Facial Visibility, 7 Storytelling, 8 Audio Clarity, 9 Subtitle Accuracy, 10 Composition, 11 Pacing, 12 Retention, 13 Visual Effects, so that higher priorities always win over lower ones.

  • Trigger/input: any conflict between competing editing goals.
  • Observable result: the higher-priority goal prevails.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a decision that violates the order is corrected.
  • Continuation: the ordered decision set is carried into the edited clip.
  • Owner: Clip Editor.

FR-27 — Quality control with REVISE (explicit) As a Post-production Reviewer / Quality Control, I should verify that original voice is present, voice is natural, subtitles match speech, speaker is correct, face is uncropped, eyes and mouth are visible, subtitles do not cover faces, the Golden Moment is genuinely interesting, context is sufficient, the ending is natural, audio is clear, music is not too loud, zoom has a reason, editing matches topic and emotion, and no unnecessary effects exist — and on failure trigger REVISE — so that no failing clip is accepted.

  • Trigger/input: a produced clip.
  • Observable result: a pass, or a REVISE that returns the clip to editing with the failing rule recorded.
  • Access state: protected, Post-production Reviewer / Quality Control and Creator / Video Owner.
  • Failure/recovery: a failing clip is never accepted; REVISE is the recovery path.
  • Continuation: a passing clip proceeds to Exports; a revised clip returns to Review after re-editing.
  • Owner: Review.

FR-28 — Absolute rules enforcement (explicit) As a Creator / Video Owner, I should have the 29 absolute rules enforced throughout: never remove original voice; never replace a human voice; never create AI voice; never change dialogue; never fabricate dialogue; never fabricate subtitles; never change meaning; never cut context misleadingly; never create fake hooks; never select Golden Moments by volume alone; never force clip count; never use the same template for all videos; never zoom randomly; never add random effects; never cover faces with subtitles; never crop eyes; never crop mouth; never create unstable framing; never remove emotional pauses; never remove important reactions; never remove natural laughter; never remove voice character; never make editing too busy; never prioritize effects over content; never produce clips just to fill quota; never claim a render is finished before it is; never use fake progress; never produce empty files; never produce broken video.

  • Trigger/input: every stage of processing.
  • Observable result: no output violates any of the 29 rules.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: a violation blocks the output and is reported with the failing rule.
  • Continuation: only compliant output proceeds.
  • Owner: Clip Editor and Review (with Analysis, Golden Moments, Exports).

FR-29 — Original voice preservation and limited enhancement (explicit) As a Creator / Video Owner, I should have the original speaker voice preserved — voice character, intonation, accent, articulation, expression, laughter, natural breathing, reactions, word emphasis, speech rhythm, vocal emotion, and vocal character — with audio enhancement limited to noise reduction, hum removal, hiss reduction, EQ, compression, de-essing, loudness normalization, limiter, and intelligibility improvement, and no over-processing when the voice is already good.

  • Trigger/input: the clip's original audio.
  • Observable result: the original voice intact with only the permitted enhancements applied.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: any processing outside the permitted list, or over-processing of an already-good voice, is reverted.
  • Continuation: the preserved voice is carried into the edited clip.
  • Owner: Clip Editor.

FR-30 — Source video prerequisite (required_inference) As a Creator / Video Owner, I should have the source video uploaded before analysis, ranking, editing, review, and export can continue, so that every downstream stage operates on a durable source.

  • Trigger/input: a selected long-form source video.
  • Observable result: the source video durably stored and downstream stages unlocked.
  • Access state: protected, Creator / Video Owner.
  • Failure/recovery: an upload failure is reported with the failing condition and the source can be re-selected and retried.
  • Continuation: analysis begins on the stored source.
  • Owner: Upload.

FR-31 — QC gate before final export (required_inference) As a Post-production Reviewer / Quality Control, I should have a clip pass quality-control review before final export, so that no failing clip reaches a ready-to-publish MP4.

  • Trigger/input: a produced clip.
  • Observable result: only QC-passed clips are available for export.
  • Access state: protected, Post-production Reviewer / Quality Control and Creator / Video Owner.
  • Failure/recovery: a failing clip is returned via REVISE and cannot be exported.
  • Continuation: a passing clip proceeds to Exports.
  • Owner: Review.

FR-32 — Durable backend execution (required_inference) As a Creator / Video Owner, I should have durable uploads, chunked analysis, candidate ranking, editing, review state, revisions, and exports executed and persisted by the backend, so that work survives across sessions and resumes correctly.

  • Trigger/input: participant actions on the protected surfaces.
  • Observable result: durable state that resumes at the correct point on return.
  • Access state: protected, Creator / Video Owner and Post-production Reviewer / Quality Control.
  • Failure/recovery: a backend failure is reported with the failing condition; completed state remains intact.
  • Continuation: the participant resumes from the last durable state.
  • Owner: system_process (supporting the first-party surfaces).

FR-33 — Self-service enrollment and returning verification (required_inference) As a Creator / Video Owner, I should establish my identity self-service on first use and verify it on return, so that my source videos, clip edits, reviews, and exports remain bound to me and resumable.

  • Trigger/input: first use or return.
  • Observable result: an established identity on first use; verified identity on return.
  • Access state: anonymous entry on Landing, Login, and Sign Up; protected thereafter.
  • Failure/recovery: an enrollment or verification failure is reported with the failing condition and can be retried.
  • Continuation: the participant enters or resumes protected work.
  • Owner: Sign Up and Login.

FR-34 — Reviewer invitation or provisioning (required_inference) As a Post-production Reviewer / Quality Control, I should reach protected review work through invitation or provisioning, so that review is bound to the correct participant.

  • Trigger/input: an invitation or provisioning for the reviewer role.
  • Observable result: protected access to Review.
  • Access state: anonymous entry on Login; protected thereafter.
  • Failure/recovery: a failed verification is reported with the failing condition and can be retried.
  • Continuation: the reviewer enters Review and acts on awaiting clips.
  • Owner: Login.
Page 10 of 22

4. User Personas

Page 11 of 22

Creator / Video Owner

Product context: The person who supplies a long-form video — podcast, interview, talk, or vlog — and needs it turned into ready-to-publish 9:16 short-form MP4 clips. They are not a video engineer; they work desktop- and laptop-first, often in a dark room, and they judge the tool by whether the output feels professionally edited and is safe to publish.

Primary goal: Get professional short-form clips from a long video automatically, with the original voice intact, accurate subtitles, face-focused framing, and a result that feels edited by a professional editor rather than merely cut by AI.

Distinct accepted responsibilities: Submitting the long-form source video; observing progressive chunked analysis; reviewing the globally ranked Golden Moment candidates and their Golden Score breakdowns; reviewing the edited clip's subtitles, framing, zoom reasons, silence classification, audio mix, and color; judging whether the output is safe to publish; exporting approved clips as ready-to-publish MP4.

Relevant inputs or decisions: Which source video to submit; which ranked candidates proceed to editing; whether the edited clip's choices (subtitle style, framing, zoom, silence handling, audio mix, color) are acceptable; whether to export.

Interactions with other accepted participants: The Creator / Video Owner hands produced clips to the Post-production Reviewer / Quality Control for review, and receives revised clips back when a QC item fails. The Creator / Video Owner can also act on Review.

Observable success: A ready-to-publish MP4 at 1080×1920, 9:16, with the original voice intact, subtitles matching speech, faces uncropped with eyes and mouth visible, subtitles not covering faces, a genuinely interesting Golden Moment, sufficient context, a natural ending, clear audio, justified zoom, and no unnecessary effects.

Source-backed constraints: The Creator / Video Owner's work is bounded by the 29 absolute rules and the 13-level priority order; nothing in the pipeline may remove or replace the original voice, change or fabricate dialogue, subtitles, or hooks, change meaning, cut context misleadingly, or produce clips merely to fill a quota.

Page 12 of 22

Post-production Reviewer / Quality Control

Product context: The role that checks each produced clip against the quality-control rules before it can be exported. They work on the same protected surfaces as the Creator / Video Owner but their responsibility is verification, not authoring.

Primary goal: Ensure no failing clip is accepted, and that every exported clip genuinely meets the quality-control bar.

Distinct accepted responsibilities: Verifying that original voice is present and natural; that subtitles match speech; that speaker attribution is correct; that faces are uncropped with eyes and mouth visible; that subtitles do not cover faces; that the Golden Moment is genuinely interesting; that context is sufficient; that the ending is natural; that audio is clear and music is not too loud; that zoom has a reason; that editing matches topic and emotion; and that no unnecessary effects exist. On failure, triggering REVISE rather than accepting the clip.

Relevant inputs or decisions: The produced clip and its QC checklist results; whether each QC item passes; whether to pass the clip or trigger REVISE.

Interactions with other accepted participants: The Post-production Reviewer / Quality Control receives produced clips from the Creator / Video Owner's editing work and returns failing clips via REVISE; passing clips proceed to Exports.

Observable success: Every clip that reaches Exports has passed all QC items, and every failing clip has been returned with the failing rule recorded.

Source-backed constraints: The reviewer's checks are exactly the quality-control list; a failure must produce REVISE, never acceptance. The reviewer's work is bound by the same 29 absolute rules.

Page 13 of 22

5. Core User Flows

Flow 1 — Creator / Video Owner: first use and source submission

  1. The Creator / Video Owner opens Landing anonymously and sees the instrument itself: the oversized stepped waveform with the Golden Moment region as a solid amber block, the stacked Silkscreen headline "EVERY CUT HAS A REASON" / "NO CUT IS INVENTED", the amber-bordered CTA "DROP A LONG VIDEO", and the inspector strip showing the Golden Score pixel-block stack and the mint "ORIGINAL VOICE: INTACT" chip.
  2. The Creator / Video Owner activates the CTA and moves to Sign Up.
  3. On Sign Up, the Creator / Video Owner establishes their identity self-service. On success they enter protected work at Upload. On failure, the failing condition is reported, entered values are preserved, and they can retry or move to Login.
  4. On Upload, the Creator / Video Owner selects a long-form source video. The duration is shown in tabular Space Mono. The submit control is disabled with a muted explanation until a source is selected.
  5. The Creator / Video Owner submits the source video. Upload progress advances in 2px increments reflecting real transfer — never fake progress.
  6. On success, the source video is durably stored and the Creator / Video Owner continues to Analysis. On failure, the failing condition is reported and the source can be re-selected and retried without losing the selection.

Flow 2 — Creator / Video Owner: progressive chunked analysis

  1. On Analysis, the Creator / Video Owner sees the chunk-analysis grid: each chunk is a square tile that flips from hairline outline to filled surface as it completes, carrying tiny pixel glyphs for speech, speaker, face, emotion, topic, keyword, hook, reaction, and Golden Moment.
  2. The live waveform's stepped columns breathe with the real audio level. The timeline strip runs across the top of the workspace with a ruled gutter of timecodes beside the transcript/analysis column.
  3. The Creator / Video Owner inspects a completed chunk's findings and the transcript, which stores start_time, end_time, speaker, text, confidence, sentence, and word timestamps when available, and reflects understood sentences, words, fillers, pauses, emphasis, questions, answers, punchlines, and keywords.
  4. Long videos are never forced into RAM at once; chunking, streaming, and progressive analysis are visible as the grid fills.
  5. If a chunk fails, the failure is reported on its tile with the failing condition, completed chunks remain intact, and the failed chunk can be re-run without discarding completed work.
  6. When all chunks complete, the Creator / Video Owner continues to Golden Moments.
Page 14 of 22

Flow 3 — Creator / Video Owner: Golden Moment review and global ranking

  1. On Golden Moments, the Creator / Video Owner sees all candidates merged from every chunk and globally ranked across the entire source video — never only early sections.
  2. Each candidate's Golden Score renders as a stack of discrete pixel blocks — Hook Value 20%, Emotional Impact 20%, Entertainment 15%, Information Value 15%, Curiosity 10%, Story Completeness 10%, Shareability 10% — filling one by one left to right, with the total in tabular Space Mono. Supporting considerations are shown per candidate: Context Integrity, Speaker Clarity, Audio Quality, Visual Quality, Uniqueness, Standalone Value, Audience Relevance, Ending Strength.
  3. The global-ranking pass visibly re-sorts the tiles at the end.
  4. The Creator / Video Owner inspects a candidate's score breakdown, its supporting considerations, and its source region on the timeline strip, and excludes candidates that do not meet the bar.
  5. Clip count is not fixed: it follows video length, number of Golden Moments, candidate quality, topic variety, storytelling, redundancy, and short-form potential. If only 3 candidates are good, 3 clips are produced; if 12 are good, 12 may be produced. No bad clip is made to fill a quota.
  6. Where candidates are near-identical, duplicate control retains the one with stronger hook, more complete context, better delivery, higher emotion, better visuals, and stronger ending.
  7. If no candidate survives ranking, the surface states that no Golden Moment met the bar rather than producing clips to fill a quota.
  8. The Creator / Video Owner selects which ranked candidates proceed to editing and continues to Clip Editor.

Flow 4 — Creator / Video Owner: professional clip editing

  1. On Clip Editor, the Creator / Video Owner sees the clip's timeline strip with tick marks and a playhead advancing in discrete ticks, the face-tracking overlay drawn as 2px pixel rectangles with corner ticks (separate boxes for face, eyes, mouth, shoulders), the "SUBTITLE ZONE" band provably below the mouth box, and the zoom-reason callout printing the exact percentage (100% / 106% / 112%) with its justification.
  2. The Creator / Video Owner inspects the subtitle style chosen for the clip's content type — Educational clean/modern with keyword emphasis; Comedy punchline timing and reaction-aware; Emotional minimal and clean; Storytelling natural to story rhythm; Debate strong statement and speaker clarity; Podcast clean/premium face-focused; Motivation strong keyword and clean typography — and confirms no single template is reused across videos.
  3. The Creator / Video Owner inspects the silence classification: unnecessary silence may be trimmed, meaningful silence is preserved, and pauses that build emotion are never removed.
  4. The Creator / Video Owner inspects the zoom reason and exact percentage, confirming the default is 100% and that Important 102–106%, Strong statement 105–110%, Emotional peak 108–112%, and Punchline 108–115% are used only when truly appropriate.
  5. The Creator / Video Owner inspects the audio mix priority — VOICE > IMPORTANT REACTION > AMBIENCE > MUSIC > SFX — with voice always most important, voice louder than music when music is used, music auto-ducking during important dialogue, and no music used if it does not help.
  6. The Creator / Video Owner inspects the color correction for exposure, white balance, contrast, skin tone, saturation, and sharpness, confirming natural results with no over-filtering.
  7. The Creator / Video Owner confirms the clip's boundaries never cut mid-word, mid-sentence, punchline, answer, important reaction, or emotional payoff, that the clip starts slightly before an important statement when context requires, and that it ends after the statement completes for a natural ending.
  8. The Creator / Video Owner confirms the hook comes from original dialogue only, with no invented hooks, no fabricated sentences, and no meaning changes.
  9. The Creator / Video Owner confirms the duration follows the content — generally 15–30 seconds, 30–60 seconds, or 60–90 seconds — and that a 22-second video was not forced into 60 seconds.
  10. The Creator / Video Owner confirms the original voice is intact — voice character, intonation, accent, articulation, expression, laughter, natural breathing, reactions, word emphasis, speech rhythm, vocal emotion, and vocal character — with only permitted enhancements applied and no over-processing.
  11. The Creator / Video Owner approves the edited clip and sends it to Review. If an editing failure occurs, the failing condition is reported and the clip remains editable without losing the source video or the ranked candidates.
Page 15 of 22

Flow 5 — Post-production Reviewer / Quality Control: review and REVISE

  1. The Post-production Reviewer / Quality Control reaches protected work through invitation or provisioning, verifies identity on Login, and enters Review.
  2. On Review, the reviewer sees the QC checklist resolving one by one as each check completes: original voice present, natural voice, subtitles matching speech, correct speaker, uncropped face, visible eyes and mouth, subtitles not covering faces, genuinely interesting Golden Moment, sufficient context, natural ending, clear audio, music not too loud, justified zoom, editing matching topic and emotion, and no unnecessary effects.
  3. The reviewer inspects each QC item's evidence on the clip, including the clip's Golden Score and supporting considerations, its speaker attribution, and its framing with the face-tracking overlay and SUBTITLE ZONE band.
  4. If every QC item passes, the reviewer passes the clip and it proceeds to Exports.
  5. If any QC item fails, the surface shows which rule failed and why, and the reviewer triggers REVISE. The clip returns to editing with the failing rule recorded, and can be re-reviewed after revision.
  6. If no clip is awaiting review, the surface states that no clip has reached review.

Flow 6 — Creator / Video Owner: export of approved clips

  1. On Exports, the Creator / Video Owner sees the approved clips that have passed review.
  2. The Creator / Video Owner exports an approved clip. The render bar advances in 2px increments reflecting real progress; no completion state is shown before the export actually finishes, and progress is never faked.
  3. On success, the finished MP4 is available at 1080×1920, 9:16, verified non-empty and non-broken, and can be downloaded as ready-to-publish content.
  4. If the export fails, the failing condition is reported and the approved clip remains available; the export can be retried. An empty or broken file is never presented as a finished export.
  5. If no clip has passed review, the surface states that no approved clip is available.

Flow 7 — Creator / Video Owner: returning to resume work

  1. The Creator / Video Owner returns and verifies identity on Login.
  2. On success, protected work resumes at their durable state — the source video, chunked analysis results, ranked candidates, clip edits, review state, revisions, and exports are all preserved.
  3. On failure, the failing condition is reported without revealing which field failed, fields remain editable, and the participant can retry or move to Sign Up.
Page 16 of 22

6. Visuals Colors and Theme

Muse: Susan Kare. Headline: Charming precision after Susan Kare — a pixel-grid instrument panel for the invisible craft of editing.

The register is calm authority plus quiet delight: a precise instrument that thinks like a senior editor and earns trust by showing its reasoning (Golden Score, speaker, silence class, zoom reason) rather than hiding it behind AI magic. The audience is desktop- and laptop-first, works in dark rooms, and already loves classic software craft.

Colour tokens (dark mode):

RoleHexUse
Background#141210Warm charcoal ground
Surface#1E1B17Slightly lifted panel surface
Border#2C2822Hairline 1px borders; 1px inset highlight
Text#F3EFE6Warm off-white type
Primary#FF8A1EAmber — the single hot signal colour: Golden Score numerals, active chunk cursor, the playhead, the one primary CTA
Accent#3FD9A4Mint — strictly data: verified/QC pass, speaker-accuracy ticks, waveform positive peaks, preserved meaningful-silence markers
Muted#8C8375Metadata, timecodes, disabled states

Colour is code, never decoration: amber = decision/moment, mint = verified/measured, muted = context. Roughly 70% charcoal, 20% warm off-white type, 8% amber, 2% mint.

Typography:

  • Headings: Silkscreen 400/700, all-caps, 0.02em tracking — display scale for section titles, 12–14px for pixel labels and chips. Never used for running copy.
  • Body/UI: Space Mono 400; 700 for numbers, timecodes, and Golden Score values. Tabular numerals everywhere; uppercase for labels, sentence case for prose. No italics; emphasis comes from weight and colour.
  • Scale: 1.25 modular on a 4pt grid — 12 / 14 / 16 / 20 / 25 / 31 / 39 / 49 / 61 / 76 / 96px. Display headline clamps 40px (375px viewport) to 96px (1280px+); section titles 25–39px; body 16px with 1.6 line-height; metadata 12–14px. Numbers always tabular.

Shape language: Hard 90° corners everywhere; 1px hairline borders; 4px pixel-step corners on chips and buttons (a single notch, never a radius). Bars, blocks, and ticks are the vocabulary: Golden Score as a stack of discrete pixel blocks rather than a smooth percentage bar; waveform as stepped columns; timeline as a ruled strip with tick marks. No soft shadows, no glass, no gradients — depth comes only from a 1px inset highlight (#2C2822) and a solid 2px offset border on the active element.

Layout: Fixed left instrument rail (72px) of pixel icons — Ingest, Chunks, Transcript, Speakers, Golden, Edit, Export — with the active icon filled amber and its label revealed on hover. Main area is a strict 12-column grid on a visible 8px baseline: a full-width timeline strip across the top of the workspace, a two-thirds transcript/analysis column with a ruled gutter of timecodes, a one-third inspector column of stacked panels (Speaker, Emotion, Silence Class, Zoom Reason, QC). Every row is an aligned label/value pair with a hairline rule. Sections are separated by a 4px amber band carrying the section name in Silkscreen. Mobile: rail collapses to a bottom bar of the same pixel icons; inspector panels stack under the transcript; timeline becomes a horizontally scrollable strip with snap-to-moment ticks.

Imagery: No photography and no illustration for atmosphere. The imagery is the interface itself: stepped waveform columns, chunk block maps, face-tracking bounding boxes drawn as 2px pixel rectangles with corner ticks, speaker chips, Golden Score block stacks, silence-class markers (trimmed vs preserved), zoom-reason callouts showing 100% / 106% / 112% as pixel numerals. The single decorative artefact allowed is one hand-drawn pixel icon set in Kare's manner (scissors, speaker, eye, spark, shield) at 16/24/32px.

Avoid: Blue–indigo primary or accent on a white ground; no #0057FF / #2563EB / #6366F1 family anywhere. Inter, Roboto, Arial, Helvetica, Open Sans, Lato, Poppins, or system-ui for headings or body. Gradient-blob heroes, glassmorphism, frosted panels, soft drop shadows. A grid of identical hover-lift cards; cards only as aligned ruled panels with no lift. Smooth cinematic easing, floating particles, bouncy micro-interactions, decorative motion. Fake progress bars or a "render complete" state shown before the export actually finishes. Any visual that implies AI voice replacement, dubbing, or a changed dialogue line. Photography or stock people standing in for the product's real analysis output.

Page 17 of 22

7. Signature Design Concept

The stepped-waveform hero. The first screen is not a centred SaaS hero — it is the instrument itself. A full-bleed dark charcoal console (#141210) whose top two-thirds is occupied by a single oversized stepped waveform: 9 columns deep at desktop, spanning the full viewport width edge to edge, its peaks in warm off-white (#F3EFE6) with the Golden Moment region rendered as a solid amber block (#FF8A1E) and everything else dimmed to muted (#8C8375).

Over the amber block, flush left, a Silkscreen headline in two stacked lines — "EVERY CUT HAS A REASON" / "NO CUT IS INVENTED" — at clamp(40px, 7vw, 96px), with a single amber-bordered CTA pinned directly beneath it reading "DROP A LONG VIDEO".

To the right of the waveform, a narrow vertical inspector strip shows a live Golden Score readout (20/20/15/15/10/10/10 as pixel blocks) and one mint "ORIGINAL VOICE: INTACT" chip.

No gradient, no blob, no floating card grid — one dominant data object, one hot accent, and the product's own output as the hero image. The concept recomposes only accepted content, states, and controls: the waveform, the Golden Moment region, the Golden Score components, the original-voice integrity state, and the CTA that begins the accepted journey.

Page 18 of 22

8. Interaction Model & Motion Direction

Interaction Model: Static Motion Tempo: restrained Hero Dimensionality: flat

Landing Hero Motion Brief

  • Focal subject: the oversized stepped waveform — 9 columns deep at desktop, spanning the full viewport width edge to edge — with the Golden Moment region as a solid amber block and the remainder dimmed to muted.
  • Input → transformation → outcome thesis: the source video's own audio level drives the waveform's stepped columns; the Golden Moment region resolves into a solid amber block as the analysis identifies it; the outcome is the product's own analysis presented as the first image, with the Golden Score pixel-block stack and the mint "ORIGINAL VOICE: INTACT" chip confirming what was measured.
  • Motion vocabulary: functional and frame-accurate — 120–160ms steps with a 6-step linear easing that reads as pixel stepping, never smooth ease-in-out. The playhead advances in discrete ticks; Golden Score blocks fill one by one left to right; chunk analysis blocks flip from outline to filled as each chunk completes; the render bar advances in 2px increments. Hover states are instant colour swaps (no lift, no scale). One purposeful ambient loop only: the live waveform's stepped columns breathe with the real audio level.
  • Composed first frame: full-bleed charcoal console; the stepped waveform occupying the top two-thirds with the amber Golden Moment block; the stacked Silkscreen headline flush left over the amber block; the amber-bordered CTA pinned beneath it; the narrow vertical inspector strip to the right with the Golden Score pixel-block stack and the mint original-voice chip.
  • Reduced-motion state: the waveform renders as a static stepped arrangement with the amber Golden Moment block, the Golden Score blocks shown fully filled, and the headline, CTA, and inspector strip all whole and readable; no ambient breathing loop.
Page 19 of 22

9. Non-Functional Requirements

NFR-01 — Memory-bounded processing (explicit) Long videos must not be forced into RAM at once. Chunking, streaming, and progressive analysis are required, with global ranking performed only after all chunks complete. Rationale: explicit source constraint; also the only way long-form sources can be processed reliably.

NFR-02 — Truthful progress and completion (explicit) A render must never be claimed finished before it is finished; fake progress is forbidden; empty files and broken video must never be produced. Rationale: explicit source constraint; the product's trustworthiness depends on it.

NFR-03 — Original voice integrity (explicit) The original speaker voice must never be removed or replaced. AI voice, voice cloning, dubbing, and robot voice are forbidden. Audio enhancement is limited to noise reduction, hum removal, hiss reduction, EQ, compression, de-essing, loudness normalization, limiter, and intelligibility improvement; if the voice is already good, do not over-process. Rationale: explicit source constraint and the product's defining promise.

NFR-04 — Output format (explicit) Output default is 1080×1920, 9:16, as ready-to-publish MP4. Rationale: explicit source constraint.

NFR-05 — Subtitle legibility limits (explicit) Subtitles must be mobile-first, max 2 lines, about 2–6 words per chunk, synchronized, following speech rhythm, and must not cover faces. Rationale: explicit source constraint.

NFR-06 — Zoom limits (explicit) Dynamic zoom defaults to 100%; Important 102–106%; Strong statement 105–110%; Emotional peak 108–112%; Punchline 108–115% only when truly appropriate; otherwise 100%. No random zoom. Rationale: explicit source constraint.

NFR-07 — Audio mix priority (explicit) VOICE > IMPORTANT REACTION > AMBIENCE > MUSIC > SFX. Voice is always loudest; music auto-ducks during important dialogue; music is not used if it does not help. Rationale: explicit source constraint.

NFR-08 — Natural color (explicit) Color correction must stay natural with no over-filtering. Rationale: explicit source constraint.

NFR-09 — Justified decisions (explicit) Every editing decision must have a reason; unnecessary changes must not be made. Rationale: explicit source constraint; it is what separates professional editing from mechanical AI output.

NFR-10 — Readable text and controls at every viewport (direction) Headlines, wordmarks, labels, numbers, cards' text, and controls stay entirely inside the viewport and their container at 375px, 768px, and 1280px, wrapping or scaling (for example font-size: clamp(...) with its mobile size) to fit, and no other element covers any part of them. Imagery, decoration, and motion may be cropped, bled off an edge, overlapped, or cut as the direction asks, as long as they cover no readable text or control. Moving and scrollable content may cross the viewport or container edge by design and is judged by whether it actually moves or scrolls and whether every item becomes fully readable as it passes. With prefers-reduced-motion, a usable static arrangement is provided: items wrap into rows or allow horizontal scrolling so each item can be brought fully into view.

NFR-11 — Accessible contrast and focus (default — not specified by user) Text and controls meet accessible contrast against the charcoal ground, and focus states are visible using the 2px offset border on the active element.

Page 20 of 22

10. Tech Stack

  • Frontend: React (web application; desktop- and laptop-first, responsive down to 375px).
  • Backend: Python / FastAPI.
  • Storage: durable storage for source videos, chunked analysis results, transcripts, ranked candidates, clip edits, review state, revisions, and exported MP4 files.
  • Containerization: Docker / docker-compose.
  • Orchestration: Kubernetes only when deployment requires it.

No other technology choices are specified by the source.

Page 21 of 22

11. Assumptions and Constraints

Assumptions

  • A-01 (required_inference) — The Creator / Video Owner independently begins using NEXUS and needs a self-service first-use enrollment path; the Post-production Reviewer / Quality Control reaches protected review work through invitation or provisioning.
  • A-02 (required_inference) — Durable source videos, chunked analysis results, ranked candidates, clip edits, review state, revisions, and exports must remain bound to the correct participant and be resumable, so application-owned identity is required for the protected destinations.
  • A-03 (required_inference) — The source video must be uploaded before analysis, ranking, editing, review, and export can continue.
  • A-04 (required_inference) — A clip must pass quality-control review before final export.
  • A-05 (required_inference) — Backend execution supports durable uploads, chunked analysis, candidate ranking, editing, review state, revisions, and exports; the human interaction for each accepted capability remains on its first-party surface.
  • A-06 (default — not specified by user) — The application is delivered as a first-party web application with a fixed left instrument rail on desktop that collapses to a bottom bar on mobile.

Constraints

  • C-01 (explicit) — Original speaker voice must never be removed; preserve original voice, voice character, intonation, accent, articulation, expression, laughter, natural breathing, reactions, word emphasis, speech rhythm, vocal emotion, and vocal character.
  • C-02 (explicit) — AI voice, voice cloning, dubbing, and robot voice are forbidden; audio enhancement is limited to noise reduction, hum removal, hiss reduction, EQ, compression, de-essing, loudness normalization, limiter, and intelligibility improvement; if the voice is already good, do not over-process.
  • C-03 (explicit) — Subtitles must come from original speech and must not be invented; hooks must come from original dialogue with no fabricated sentences and no meaning changes.
  • C-04 (explicit) — Never cut mid-word, mid-sentence, punchline, answer, important reaction, or emotional payoff.
  • C-05 (explicit) — Never select Golden Moments based on loud audio, moving faces, keywords, or volume alone.
  • C-06 (explicit) — Never force a fixed clip count or produce clips merely to fill a quota.
  • C-07 (explicit) — Never use one identical subtitle template for all videos.
  • C-08 (explicit) — Never zoom randomly and never add random effects.
  • C-09 (explicit) — Never cover faces with subtitles; never crop eyes or mouth; never create unstable framing or framing jitter.
  • C-10 (explicit) — Never remove emotional pauses, important reactions, or natural laughter.
  • C-11 (explicit) — Never make editing too busy or prioritize effects over content.
  • C-12 (explicit) — Never claim a render is finished before it is finished; never use fake progress; never produce empty files; never produce broken video.
  • C-13 (explicit) — Output default is 1080×1920, 9:16.
  • C-14 (explicit) — Dynamic zoom limits: default 100%, Important 102–106%, Strong statement 105–110%, Emotional peak 108–112%, Punchline 108–115% only when truly appropriate, otherwise 100%.
  • C-15 (explicit) — Subtitle limits: max 2 lines, about 2–6 words per chunk, mobile-first, must not cover faces.
  • C-16 (explicit) — Audio mixing priority VOICE > IMPORTANT REACTION > AMBIENCE > MUSIC > SFX; voice always loudest; music auto-ducks during important dialogue; do not use music if it does not help.
  • C-17 (explicit) — Color must stay natural with no over-filtering.
  • C-18 (explicit) — Long videos must not be loaded into RAM all at once; chunking, streaming, and progressive analysis are required, with global ranking after all chunks complete.
  • C-19 (explicit) — Editing decisions must be justified; unnecessary changes must not be made.
  • C-20 (explicit) — The 29 absolute rules are binding at every stage: never remove original voice; never replace a human voice; never create AI voice; never change dialogue; never fabricate dialogue; never fabricate subtitles; never change meaning; never cut context misleadingly; never create fake hooks; never select Golden Moments by volume alone; never force clip count; never use the same template for all videos; never zoom randomly; never add random effects; never cover faces with subtitles; never crop eyes; never crop mouth; never create unstable framing; never remove emotional pauses; never remove important reactions; never remove natural laughter; never remove voice character; never make editing too busy; never prioritize effects over content; never produce clips just to fill quota; never claim a render is finished before it is; never use fake progress; never produce empty files; never produce broken video.
  • C-21 (explicit) — The priority order is binding: 1 Meaning, 2 Original Voice, 3 Speaker Accuracy, 4 Context, 5 Golden Moment Quality, 6 Facial Visibility, 7 Storytelling, 8 Audio Clarity, 9 Subtitle Accuracy, 10 Composition, 11 Pacing, 12 Retention, 13 Visual Effects.
Page 22 of 22

12. Glossary

  • NEXUS — The professional AI video editor and Golden Moment Director delivered by smart-editor.
  • Golden Moment — A part of the video with high value for short-form content, judged on strong statements, surprising facts, stories, punchlines, conflict, controversial opinions, emotional reactions, confessions, funny moments, inspiration, solutions, twists, revelations, lessons, and shareable statements — never on loud audio, moving faces, keywords, or volume alone.
  • Golden Score — The weighted candidate score: Hook Value 20%, Emotional Impact 20%, Entertainment 15%, Information Value 15%, Curiosity 10%, Story Completeness 10%, Shareability 10%, plus Context Integrity, Speaker Clarity, Audio Quality, Visual Quality, Uniqueness, Standalone Value, Audience Relevance, and Ending Strength.
  • Chunk — A bounded segment of the source video analyzed progressively; each chunk is analyzed for speech, speaker, face, emotion, topic, keyword, hook, reaction, and Golden Moment.
  • Global ranking — The pass performed after all chunks complete, merging all candidates and ranking them across the entire source video.
  • Active speaker — The speaker currently talking, determined by face tracking and speaker diarization, used to focus framing.
  • Virtual camera — The mechanism used for pan, tilt, zoom, follow speaker, reframe, center face, and headroom maintenance in auto reframe.
  • Unnecessary silence — Silence that may be trimmed.
  • Meaningful silence — Silence that must be preserved, including pauses that build emotion.
  • Adaptive subtitle — A subtitle style chosen per content type (Educational, Comedy, Emotional, Storytelling, Debate, Podcast, Motivation) rather than one template for all videos.
  • Hook — The opening drawn verbatim from original dialogue; never invented, never a fabricated sentence, never a meaning change.
  • Clip boundary — The start and end of a clip, which must never cut mid-word, mid-sentence, punchline, answer, important reaction, or emotional payoff.
  • Duplicate control — The resolution of near-identical clips by retaining the one with stronger hook, more complete context, better delivery, higher emotion, better visuals, and stronger ending.
  • QC — Quality control: the checklist verifying original voice present, natural voice, subtitles matching speech, correct speaker, uncropped face, visible eyes and mouth, subtitles not covering faces, genuinely interesting Golden Moment, sufficient context, natural ending, clear audio, music not too loud, justified zoom, editing matching topic and emotion, and no unnecessary effects; on failure, REVISE.
  • REVISE — The required action when a clip fails quality control; the clip returns to editing with the failing rule recorded.
  • Ready-to-publish MP4 — The final export: MP4 at 1080×1920, 9:16, verified non-empty and non-broken.

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Landing: Open instrument hero anonymously
Sign Up: Create owner identity
Upload: Select long-form source
Upload: Submit source video
Analysis: 1. Observe progressive chunk analysis
Analysis: Inspect chunk findings and transcript
Analysis: 2. Re-run failed chunk
Golden Moments: Review globally ranked candidates
Golden Moments: Inspect Golden Score breakdown
Golden Moments: Exclude weak duplicate candidate
Golden Moments: Select candidates for editing
Clip Editor: Inspect subtitle style and tracking
Clip Editor: Inspect silence class and zoom reason
Clip Editor: Inspect audio mix and color
Clip Editor: Approve clip for review
Review: Check clip QC evidence
Clip Editor: Correct failing rule
Review: Pass revised clip
Exports: Export approved clip
Exports: Download finished MP4
Login: Verify identity on return
Upload: Reselect source after failure
Login: Navigate to Sign Up
Sign Up: Retry enrollment
Login: Retry verification after failure

No completed page designs yet.

Completed design pages will appear here when they are ready to preview.

Landing: Open instrument hero anonymously
Sign Up: Create owner identity
Upload: Select long-form source
Upload: Submit source video
Analysis: 1. Observe progressive chunk analysis
Analysis: Inspect chunk findings and transcript
Analysis: 2. Re-run failed chunk
Golden Moments: Review globally ranked candidates
Golden Moments: Inspect Golden Score breakdown
Golden Moments: Exclude weak duplicate candidate
Golden Moments: Select candidates for editing
Clip Editor: Inspect subtitle style and tracking
Clip Editor: Inspect silence class and zoom reason
Clip Editor: Inspect audio mix and color
Clip Editor: Approve clip for review
Review: Check clip QC evidence
Clip Editor: Correct failing rule
Review: Pass revised clip
Exports: Export approved clip
Exports: Download finished MP4
Login: Verify identity on return
Upload: Reselect source after failure
Login: Navigate to Sign Up
Sign Up: Retry enrollment
Login: Retry verification after failure