Product + craft audit Current system → elite human UGC

You built a strong studio OS.
The missing product is the editor’s judgment.

Spielberg safely receives messy footage, keeps context, runs an editor, supports targeted revisions, and delivers one final. But the path from technically complete to hand-montaged, native, magnetic is still mostly written advice—not a repeatable product guarantee.

888automated tests passed
0representative finished UGC masters checked in
5durable brief fields today
1complete final delivered per round
Read-onlyno product code changed
Evidence boundary

This is a static architecture, UX, runtime-contract, and market-craft audit—not a blinded score of customer output. The repository contains a synthetic review fixture and a real repeated-take research set, but no representative completed UGC masters or post-publish retention data. That absence is itself the largest proof gap.

01 · The verdict

Reliability gets a video delivered.
Judgment gets it watched.

01

A perfect montage is not more effects. It is the right 24 frames at the right moment, proof before doubt, sound carried across the cut, and every decorative move motivated.

The current system knows much of this grammar in its prompts and skills. What it does not yet do is turn those choices into required, inspectable evidence attached to the exact delivered file.

Keep

The production-minded spine

Tenant isolation, immutable raw media, resumable upload, one concise conversational reply, persistent per-montage context, idempotent delivery, and redo → fragment → approve → stitched-final semantics are unusually solid.

Change

The completion contract

Today, “done” reduces to a successful assembler run plus pieces.json. Delivery can select the newest MP4 and check ownership, duration, and dimensions without a mandatory story, audio, caption, crop, freeze, or visual-proof receipt.

Prove

The creative promise

“Perfect” needs a gold set of messy UGC packs edited by strong human social editors, blind pairwise ratings, and—in production—retention, saves, shares, clicks, and revision decisions joined back to edit choices.

The single decision

Do not spend the next cycle adding transitions.

Make the system understand the footage before it asks, formalize a story plan before it renders, and refuse delivery until an exact-master quality contract passes.

02 · Current state

A capable agent-operated toolkit,
not yet a deterministic craft pipeline.

01

Receive

Telegram or resumable web uploader stores immutable originals and basic duration/dimension metadata.

Strong primitive
02

Brief

One persistent chat extracts platform, 15/30/60s, mood, captions, and music; richer intent remains free-form.

Too shallow
03

Agent edit

One autonomous editor session is expected to apply a deep library of directing, raw-edit, motion, sound, and QA doctrine.

Prose-dependent
04

Assemble

A round becomes “done” when assembly exits successfully and a pieces manifest exists.

Weak gate
05

Deliver

One whole MP4 is sent. A targeted redo yields a fragment; approval stitches the next whole master.

Good interaction model
CURRENT assemble_round exited 0
+ pieces.json exists
TARGET master.sha matches
+ quality_contract.status = pass

03 · Gap to the best human edits

Correct → clean → coherent → native → magnetic.

This is a capability-enforcement judgment, not an empirical output score. The service is strong at correctness, has meaningful clean-up tools, and can reach higher when the agent executes the playbook. It does not yet prove those upper levels consistently.

01CorrectStrong

The right tenant, montage, round, playable format, duration, and dimensions.

02CleanPartial

Bad takes, repeats, silence, captions, mix, and technical defects handled.

03CoherentInconsistent

One promise, compressed story, early proof, earned CTA, no narrative waste.

04NativeUnproven

Feels born in the feed: creator voice, safe UI, intentional imperfection, platform rhythm.

05MagneticUnproven

The eyebrow lift, tactile click, delayed reveal, emotional hold, and surprise that earn attention.

Editorial decision Service contract today Elite human cut Mechanism to add
Moment selection Transcript, waveform, duplicate, and timing evidence; limited durable visual-performance data. Picks gaze, credibility, gesture peak, microexpression, product visibility, and the revealing half-second after the line. Performance atlas per phrase/take: emotion, eye contact, gesture, blink, energy, subject/product, technical confidence.
Hook Strong written doctrine, but no required uploaded-footage hook lab or first-frame acceptance gate. Picture, voice, text, and sound converge on one promise in the first event—not a title-card throat clear. Three real hook assemblies: result-first, contrarian/confession, sensory demo; contact sheet + 0–2.5s preview.
Story compression An agent can build beats; no required promise → proof → payoff ledger or alternate assembly. Moves proof before explanation, deletes correct-but-unnecessary lines, preserves only setup, escalation, objection, payoff. Edit-decision map with each beat’s intent, claim, proof asset, payoff, transition reason, and confidence.
Pacing Docs reject metronomic cutting; execution remains prompt-led and has no retention feedback. Pattern interrupts at narrative turns, holds when comprehension or emotion needs time, accelerates only for tension/payoff. Intentional cadence: beat role + shot-length rationale + viewer-response data linked to the timeline.
Sound Deterministic music and loudness tooling; limited productized dialogue repair, ambience continuity, or first-class split edits. J-cuts the next idea, L-cuts voice over proof, smooths room tone, preserves breath and tactile click/spray/keyboard as evidence. Five-part sound scene: dialogue, ambience, native product sound, music, sparse SFX—with independent A/V in/out.
Captions Real word timing and art direction; limited languages, fixed patterns, and a generic 9:16 safety model. Text adds the number, objection, or context voice cannot. It is not karaoke. Placement responds to face, product, UI, and device. Semantic captions + term review + face/product/UI collision tests in TikTok/Reels device previews.
Picture finish Scale/pad, LUTs, and size checks; no guaranteed semantic reframe, shot match, skin/exposure gate, or stabilization path. Corrects exposure/WB first, protects skin and product color, keyframes crop as attention moves, keeps useful phone texture. Tracked reframe + shot match with manual override and an explicit “keep for trust / remove as defect” decision.
Proof + CTA Can be requested in free text; not mandatory or verified in canonical cut artifacts. Shows a result, use, comparison, receipt, or real reaction while the claim is made; CTA is earned by the payoff. Proof ledger connecting every claim to a visible beat and checking CTA wording, visibility, and safe placement.

04 · What winning social creative teaches

The feed rewards visible proof, human specificity, and native behavior—not effect density.

The numbers below are platform-reported advertising studies and creative-center observations. They are directional craft signals, not causal guarantees for organic performance.

TikTok creative research
2s

The opening is audiovisual.

TikTok reports the first two seconds account for half of recall and awareness impact. Its guidance favors vertical, sound-on, authentic work where every visual and audio element has a reason.

Official source ↗
TikTok-reported associations
+175%

Show the product in use.

TikTok’s Creative Accelerator associates product-in-use footage with higher conversion and recommends live action, natural light, facial close-ups, a shot list, and audio.

Official source ↗
TikTok study
71%

Over-polish can reduce trust.

Respondents said brand posts should not feel too polished; TikTok’s coded ad study reports casual, UGC-friendly work nearly doubled several early-watch and engagement outcomes.

Official source ↗
Meta split-test summary
−34.5%

Native format is structural.

Meta reports lower cost per result for Reels creative built in 9:16 with audio and key elements inside safe zones versus image ads in its cited tests.

Official source ↗
Annotated ad pattern

Proof as format

TikTok’s True Classic example moves oddly satisfying hook → before/after → user validation in lived-in scenes → pack/promo CTA. Vaseline Verified built an entire repeatable format by testing community hacks, then visibly verifying or debunking them.

Current Creative Center patterns

Specific, lived-in, close

TikTok’s surfaced high-performing creative is repeatedly described through direct before/after, familiar scenes, curiosity leads, immersive close-ups, and scenario-specific customization—not generic cinematic finish.

Human editor method

Build a performance

Professional editors string out takes, mark reactions and improvs, select around emotional beats, and sculpt one consistent performance. Tiny looks and even a blink at a cut can change perceived credibility.

05 · The conversation experience

Chat should reveal that the editor saw the footage.

Today, the service counts files and duration, then asks generic brief questions. The north star is a content-grounded collaborator that names moments, flags risk, proposes a story, and makes every revision exact and reversible.

S
SpielbergFootage understood · 9 clips
Today, 17:42
Need a Reel for this serum. Women 25–40, make them buy; no medical claims.
Footage read I found 14 usable moments across 9 videos: 3 demos, 2 testimonials, 4 lifestyle shots, and 5 details. Clips 6 and 9 are noisy but visually useful. Take 3 has the strongest eye contact; the bottle drop is the best first-frame event.
I’d make a 20–25s vertical conversion Reel: result → skepticism → demo → visible proof → “Shop now.” One choice before I cut: should it feel organic/confessional or polished/confident?
Organic. Ukrainian captions. Music low under voice.
Brief locked 9:16 · 20–25s · organic/confessional · Ukrainian · Shop now · no medical claims. I’ll keep the real laugh and handheld close-up, clean the dialogue, and avoid the overexposed beauty shot.
0–2s Hook2–6s Pain6–15s Proof15–21s Objection21–24s CTA
P0 · Broken promise

The web “Chat” is not the same agent chat.

Telegram invokes the persistent montage session. The review composer only appends to an inbox relayed to the human owner, while the UI says the agent will reply. In delivered one-page mode, Chat and Comments are hidden entirely.

Fix:

One event log and one conversation gateway. A message from either surface invokes the same session and projects the same reply to both.

P0 · Wrong interpretation

A voice brief is treated as footage.

Onboarding says the brief can be spoken, but a normal voice note becomes an audio asset unless a narrow revision state is active. It is not transcribed into the conversation.

Fix:

Voice in intake/brief becomes a chat turn by default; only “use this as source audio” makes it media.

P1 · Trust gap

A revised fragment can dead-end.

The state waits for approval, but the fragment can arrive with only a label—no “approve or say what is off”—and the model context omits the pending approval target.

Fix:

Every fragment names the exact delta, locked spans, and next action. Add Undo and Compare.

P1 · Generic brief

Five fields cannot encode a winning ad.

Goal, audience, promise, offer, proof, CTA, must-use/avoid, compliance, authenticity tolerance, crop, brand, reference behavior, and variants are not durable structured state.

Fix:

A confidence- and provenance-aware creative contract; ask only the single highest-impact missing question.

PRECISE REVISION

From vague feedback to a reversible transaction

User “Hook is slow. Start with the bottle drop and lower music under speech.”

Agent “I’ll replace only 0:00–0:02 with upload 4’s bottle drop and lower music 4 dB under speech. Everything else stays locked. I’ll send the opening first.”

Fragment “Opening · 0:00–0:03. The drop now lands on the beat; music is lower under the first line. Approve or tell me what is still off. Compare V1 · Undo

User “Approve.”

Agent “Locked. V2 changes only opening 0:00–0:02 and speech/music balance; all other frames match V1.”

CURRENT UX WEAK-POINT REGISTER

Fix the promises around the edit, too.

P0
No outage-safe progression

“That’s all” and “start” depend on model interpretation. There is no deterministic done/start path, and a failed natural-language turn is not replayed despite recovery copy promising a later answer.

P0
Input can be lost on corrupt state

The user is told the message will be saved, but archiving happens only after successful registration. Save the raw event before reading mutable state.

P1
Progress is pull-only

Real editor narration exists, but there is no client-visible milestone plan, ETA range, next verified unit, cancel, or meaningful-delay receipt. Copy can imply proactive follow-up that is intentionally disabled.

P1
Memory under-populates

The architecture separates client, preference, and project knowledge, but normal chat writes mainly preferences; identity, audience, provenance, conflicts, and project indexing can remain incomplete.

P1
Uploader lacks recovery controls

Resumption is solid, but there are no thumbnails, remove/cancel, duplicate warning, per-file retry, reorder, or exact error recovery; one fatal path can leave completion stuck.

P1
Journey is not localized end to end

Chat follows the user’s language while review UI is hardcoded English, including a hardcoded operator name. Ukrainian/Russian users cross a language and ownership boundary at delivery.

P1
Opaque daily chat cap

A 40-turn default can interrupt revision with no remaining-turn indicator or reserved approval/blocker budget.

P2
No real browser journey in CI

Tests exercise contracts and source shape, not rendered mobile WebView, localization, accessibility, offline/reload, uploader failure, or the cross-surface conversation promise.

06 · Current result surface

The review page feels like an operator console,
not a creator’s result moment.

Current desktop Review Studio with large player, editing timeline, Comments, Library, and Chat tabs
Live local review surface captured during this audit1280 × 800
01

Split mental model

Telegram promises a simple conversation and one result. Desktop opens a technical timeline with Comments, Library, and Chat; mobile switches to Cuts & trims, Visual effects, and Sound. The experience changes product category mid-journey.

02

False continuity

“Vadym’s agent will reply here” hardcodes an operator identity and suggests shared conversational memory that the web path does not have.

03

Wrong first question

The result moment should lead with “Does this story feel right?” and three grounded actions—not expose an NLE-shaped surface before the creator asks for control.

Target result center

Video first. Decision second. Timeline only on demand.

  • Watch: full-frame final, real TikTok/Reels safe-zone preview, cover frame.
  • Understand: “Hook / proof / payoff / CTA” map and QA receipt in human language.
  • Act: Approve · Change this moment · Compare hooks · Export pack.
  • Chat: one compact contextual composer always visible, identical memory on every surface.

07 · Video quality and pipeline

Turn expert advice into machine-verifiable evidence.

P0Exact-master identity + QA

Never infer “final” from the newest MP4.

Bind the approved master path and SHA to the round. Require full decode, expected fps/duration/audio, black/frozen frames, word-boundary joins, A/V sync, caption text/safe zones, face/product crop, loudness/true peak, and required story beats. Any failed check repairs or blocks delivery.

P0Media intelligence

Build a box of moments before asking.

Per asset: content hash; codec/fps/VFR/HDR/rotation; audio health; shots; word-level ASR; faces, hands, product and action tracks; blur/shake/exposure; duplicate clusters; emotion/performance; native-sound events. Present contact sheets and ranked ranges, not just file count.

P0General planner

Plan by story archetype, not one talking-head path.

Manifest-driven strategies for testimonial, product demo, founder/coach, lifestyle, mini-vlog, event/travel, interview/multicam, screen recording, and mixed photo/video. Each cut records selection reason, story role, native sound, crop intent, and confidence.

P0Render contract

Close the frozen/blank-frame escape hatch.

Normalize CFR, pin one tested HyperFrames version, restore missing composition-library gates, use the proven capture path, fail on zero visual samples, and attach snapshot/motion checks to the delivered master—not only to documentation.

P1Human finish

Repair what viewers feel.

Dialogue isolation/denoise/de-reverb/EQ/de-ess, room-tone continuity, beat/phrase-aware music, tactile nat sound, sparse SFX, semantic reframe, stabilization, exposure/WB/skin/product match, multilingual term-aware captions.

P1Fast creative choice

Render decisions before pixels.

Show the creative contract, story map, three hook candidates, and a low-cost proxy first. The user can lock a concept and important moments before a full-quality render. This reduces expensive random rerolls.

quality_contract.json PASS REQUIRED
{
  "master": { "path": "round-02/master.mp4", "sha256": "…" },
  "story": { "promise": "pass", "visibleProofBy": "00:06.2", "cta": "pass" },
  "picture": { "blackFreeze": "pass", "faceProductCrop": "pass", "safeZones": "pass" },
  "sound": { "wordCuts": "pass", "dialogue": "pass", "loudnessTruePeak": "pass" },
  "captions": { "scriptDiff": "pass", "collision": "pass" },
  "status": "pass"
}

08 · Concrete output examples

Three edits that feel chosen, not generated.

01

27s product testimonial

Proof-first confession

Best for: beauty, wellness, home, apps, small consumer products.

0–2.2Result + skepticism
2.2–5Real problem
5–10Tactile demo
10–16Comparable proof
16–21Objection
21–24Reaction hold
24–27Earned CTA

Human choices: first frame already has product motion; strongest skeptical line from any take; keep the eye-roll or laugh; L-cut voice across the close demo; preserve the snap/spray/click; hold 8–12 frames after genuine relief.

Avoid: logo card, karaoke captions, every-cut whoosh, unmatched before/after, compulsory LUT, or CTA before proof.

02

30s founder / coach insight

Contrarian receipts

Best for: SaaS, creators, education, professional services.

0–2.8Receipt + claim
2.8–6Stakes
6–13Demonstration
13–17Skeptic turn
17–24Readable proof
24–28Face + conviction
28–30Value CTA

Human choices: moving screenshot in the first frame; hard cut on eye contact; J-cut the next phrase under proof; slow down for the receipt; carry clicks/keyboard as native sound; select the confident microexpression, not only clean diction.

Avoid: “stop scrolling,” counterfeit reply-to-comment UI, unreadable screen recording, or “follow for more” unless follower growth is the actual goal.

03

30s lifestyle / mini-vlog

Mess → attempt → sensory payoff

Best for: travel, food, events, fitness, transformations.

0–3Peak tease → reset
3–8Escalating setup
8–14Keep complication
14–22Process + audio bridge
22–27Long payoff
27–30Proof + loop

Human choices: keep the spill, miss, laugh, or doubt that creates stakes; non-metronomic 0.4/0.8/1.3s setup holds; cut on movement and native sound; alternate close/medium/wide; give the reveal 1.5–2.5s and include the reaction after it.

Avoid: equal clip lengths, file order, stabilizing every honest wobble, crash zooms on a timer, or looping a motion that destroys comprehension.

09 · North-star product

One chat. A visible plan. A verified result.

01Uploadimmutable originals + proxies
02Understandmedia graph + performance atlas
03Contractgoal, promise, proof, CTA
04Proposestory map + three hooks
05Proxychoose before full render
06VerifyQA, repair, exact SHA
07Delivervariants + reversible chat
One source of truth

Every surface is the same conversation.

Telegram, uploader, result page, timestamp comment, and voice all become contextual events in one montage session.

Visible reasoning

Show selections, not chain-of-thought.

“I chose take 3 for eye contact; kept clip 6 only for its close-up; proof appears by 0:06.” Users can correct facts and taste.

Purposeful variants

Different strategies, not random rerolls.

Conversion / Organic / Energy—or three hooks on one locked body—with clear differences and a fast compare path.

Reversible control

Words map to exact moments.

Frame, word, scene, upload, and subject references; unchanged spans lock; every action has a diff, Undo, and Compare.

Honest progress

Milestones with evidence.

Receiving → validating → analyzing → selecting → rough cut → captions/audio → QC → rendering, with ETA range and blocker.

Learned taste

Remember decisions, not moods alone.

Pace, shot density, caption energy, crop tolerance, nat-sound preference, zoom restraint, imperfection budget, and accepted revisions.

10 · Prioritized roadmap

Build trust first. Then taste. Then scale the moat.

Now · P0Trustworthy first cut
  1. Exact-master QA receipt: no newest-file inference; strict joins; decode, black/freeze, A/V, captions, crop, sound, story checks.
  2. Media graph + performance atlas: what, who, action, quality, emotion, proof, duplicates, native sound, usable ranges.
  3. Creative contract + story plan: goal, audience, promise, proof, CTA, must-use/avoid, compliance, authenticity.
  4. One conversation gateway: Telegram and web share state; voice brief works; deterministic start/status/retry survive model outage.
  5. Golden UGC evaluation: diverse human-edited packs, blind pairwise review, defect checks, revision and preference metrics.
Exit when

Every final is artifact-bound and QA-passed; users see a footage-grounded plan before render; the benchmark can say whether output beats a baseline.

Next · P1Human finish
  1. Dialogue repair, ambience continuity, J/L decisions, native-sound scene, beat/phrase-aware music.
  2. Face/hand/product-aware reframe; shot-match exposure/WB; skin/product protection; stabilization with intentional-imperfection flags.
  3. Semantic multilingual captions, brand/term lexicon, per-placement device safe zones and collision checks.
  4. Three hook candidates, fast proxies, exact target revisions, Undo/Compare, locked unchanged spans.
  5. Purpose-built 9:16 / 1:1 / 16:9 packages, cover frames, clean/burned captions, music/no-music when needed.
Later · P2Creative learning moat
  1. Safely freeze and analyze actual TikTok/Reels references: cadence, caption density, proof order, crop, sound, imperfection.
  2. Learn per-client and per-vertical taste from approved revisions—not a single house style.
  3. With consent, join 2s/6s hold, retention dips/spikes, completion, saves, shares, clicks, and conversion to edit decisions.
  4. Deliver an optional professional bundle: clean master, captions, covers, cue/decision sheet, rights ledger, QA receipt, XML/EDL.
MEASURE THE PRODUCT

North-star metrics

First-pass acceptance% approved without creative rerender
Meaningful revision turnsmedian to final approval
Wrong-moment rate“not that clip/word/scene” per job
QA escape ratedefects found after delivery
Time to understood planupload complete → grounded proposal
Time to proxyapproved contract → watchable decision
Human preference win rateblind pairwise vs baseline/human gold
2s / 6s / completionwith consent, by hook and beat decision

11 · Evidence trail

What this assessment is built on.

Method: traced the live upload, session, editor, assembly, delivery, review, revision, memory, and QA paths at commit cc9fc44; ran the repository’s full 888-test suite; opened the current review surface at desktop and mobile sizes; inspected engine skills, fixtures, and evaluation artifacts; benchmarked official TikTok/Meta guidance, public high-performing creative patterns, professional editing workflows, current AI editor controls, and primary technical research.

Interpretation rule: documented ability is not counted as a product guarantee unless the runtime produces durable evidence and the delivery gate requires it. Platform metrics remain attributed to the platform.

Final thought

The product should not merely say
“your montage is ready.”

It should be able to prove: “I understood what happened, chose why each moment belongs, protected what makes it human, and verified this exact file before I sent it.”

Back to top ↑