Nave
ENRIQUE the faceless machine
STEP 0 / 8
ES · Español
Nave presents — a production of the Content OS

A YouTube channel with no face, no camera,
no editor.

How a person — or a business — starts a faceless channel from zero: an idea goes in, an upload-ready video comes out. Narration, illustrated stills in one locked style, on-screen text, music, thumbnails, description. Eight stations, three decisions, one chat. This page is the map, the demo and the step by step.

8stations
3decisions you make
1chat, start to finish
0recordings
0editors
SCROLL
Cold open — the test piece, produced start to finish

Eighteen seconds. Zero human hands between the idea and the file.

A three-sentence test script, spoken by a cloned voice, cut into four shots on its own phrase boundaries, four illustrations generated in one style from one reference image, assembled and rendered with the on-screen words as real type. It is a smoke test, not a hit video — and it is exactly the pipeline a ten-minute video runs through, 150 images instead of four.

TEST PIECE · 18 s · 4 SHOTS
18.2 s · cloned voice · 543 frames rendered3 script segments → 54 timed words → 4 shots → 4 images → 1 master
style key
THE STYLE KEY — one image, attached to every generation
shot 1
shot 2
shot 3
shot 4
THE FOUR SHOTS — same paper, same character, same light
timeline contact sheet
THE TIMELINE — one frame every 2 s: images land on their phrase; on-screen words are type, not pixels
Who it's for — and what you need

One person. Or one business. Starting from zero.

Faceless explainer channels are usually run by hand: a chat prompt for ideas, another for the script, a voice tool, a transcriber, one image prompt per timestamp, a drag-and-drop edit, an upload form. It works — channels with a handful of videos and tens of thousands of subscribers exist in every explainer niche — and it burns a day per video. These eight stations keep that shape and run it from files.

PERSONA creator without a face on camera

You have a topic you can explain and no wish to film yourself. You bring a niche and a voice — cloned, or the free local one — and the machine writes, illustrates, assembles and packages.

  • a niche you can talk about for 50 videos
  • a voice (yours cloned, or a free one)
  • Claude Code + the Content OS
BUSINESSA brand that wants an owned channel

Explainers about your world — the questions customers ask before they buy — published every week under your brand, without a studio day. Same machine, your palette, your style key.

  • a brand layer (colors, voice, audience)
  • 3–6 competitor channels to model
  • the keys: voice, images, transcription
COSTWhat a ten-minute video costs

Estimated, at published prices. You pay the providers directly; the OS orchestrates and adds nothing.

  • voice ≈ 10k characters (ElevenLabs) or free (Kokoro)
  • ≈150 images at 1K ≈ $6 (KIE) or ≈150 Higgsfield credits
  • 3 thumbnails at 2K · render local, free, ~20 min
Honest line: the machine makes the video. Views come from the idea and the script — that is why the first two stations exist, and why the director decides there.
What it runs on — connected, not improvised

Claude produces start to finish. These are its hands.

The map — the machine, top to bottom

How does an idea become an uploaded video?

Eight stations. Three are gates — you decide there; the system runs the rest. Every station reads and writes files in one folder, so the line resumes wherever it stopped, and a member who already has a script or a voice file enters mid-line.

0
CHANNELname, handle, logo, banner, description from real keywords, the Studio settings — once, before the first upload
1
IDEAS3–6 competitor channels scraped; every video ranked by views ÷ channel median; 50 ideas, top 10 — you pick
2
SCRIPT2–5 donor videos beat-mapped (structure, never wording); research with sources; narration in segments — you approve
3
VOICEevery segment spoken, joined into ONE wav, transcribed word by word: the clock every picture obeys
4
STORYBOARDthe style key; the voice cut into shots on phrase ends; one scene prompt per shot; text kept as a separate field — you see the first 10
5
IMAGESNano Banana Pro, the style key attached to every generation, files named by their second, contact sheets
6
ASSEMBLYHyperFrames: images on their phrase, slow drift, on-screen type, voice at −14 LUFS with music ducked under it, chunked render, QA sheet
7
PACKAGE5 titles inside the niche molds, 3 thumbnails in the style key, description with real chapters, tags, the upload checklist
the style key
STATION 4 · the style key — the one image every later image is generated against
you decidethe system runs
Clarity — what each station really does

What does a production company do… when there is no production company?

The same jobs a research desk, a writer, a narrator, an art director, an illustrator, an editor and a publisher would do. Each one is a skill; each skill writes files the next one reads.

SYSTEM

RESEARCH DESK → ideas

Not guesses: the competitors' real view counts, ages and titles, pulled video by video. An idea is good when a channel's own audience over-reacted to it — that is a number.

YOU

EDITOR-IN-CHIEF → the pick

You choose the idea and the package direction from the top 10. Everything downstream costs money; this is the cheapest place to be wrong.

SYSTEM

WRITER → the narration

The format comes from donor videos' beat maps; the facts come from research with links; the words are written to be spoken. On-screen notes and sources never enter the voice text.

YOU

DIRECTOR → the script

You read it. Voice and 150 images follow these words, so nothing past this gate changes them. Edits are per segment — cheap.

SYSTEM

NARRATOR + ART DIRECTOR → voice and style

One WAV with word timings; one style descriptor plus one style-key image. Consistency is a reference image, not a paragraph of adjectives.

SYSTEM

ILLUSTRATOR + EDITOR + PUBLISHER → the file

Shots on phrases, images in one look, type instead of painted text, music under the voice, a verified render, titles in the molds that already work in the niche.

The reveal — the skills chain

One system. Idea in, upload out.

This is the Content OS: each station is a skill, the skills chain, and the whole line runs in one chat until the master and its package are on disk. Swap the style, the voice or the image provider and nothing else moves.

STATION 0 · once

faceless-channel-setup

  • name + handle that read as the niche
  • logo + banner in the style key, no painted text
  • description from keywords with real demand
  • Studio checklist: country, keywords, watermark, auto-dub
STATION 1 — GATE

faceless-niche-intel

  • competitors scraped video by video
  • outlier score = views ÷ channel median
  • title molds from the real corpus
  • 50 ideas scored → top 10 → you pick
STATION 2 — GATE

faceless-script

  • donor beat maps: structure, never lines
  • truth-first research, two sources per number
  • narration in segments + on-screen + claims
  • 5 alternate hooks → you approve
STATION 3

faceless-voiceover

  • cloned voice, or free local Kokoro
  • ONE wav joined as PCM (no drift)
  • word-level transcript = the clock
  • bring your own voice file: enter here
STATION 4 — GATE

faceless-storyboard

  • style preset or a style read from references
  • shots cut at phrase ends from the real timings
  • one scene prompt per shot; text as a field
  • validator refuses text-in-image → first 10 proof
STATION 5

faceless-visuals

  • Nano Banana Pro via Higgsfield or KIE
  • style key attached to every generation
  • files named by index + second
  • contact sheets, resume, regenerate rejects
STATION 6

faceless-assemble

  • HyperFrames chunks ≤ 20 s at shot starts
  • drift / push motion, on-screen type
  • voice −14 LUFS, music ducked under it
  • serial render, frame-verified, QA sheet
STATION 7

faceless-package

  • 5 titles inside the niche molds
  • 3 thumbnails in the style key, exact text
  • description with chapters from real timings
  • upload checklist → publish
RESULT

UPLOAD-READY. NEXT IDEA.

  • one folder, every file, resumable
  • the same line for video two
  • costs stated before you spend
END TO END — one chatTHREE GATES — you decide, the rest runsevery skill swappable — it's an OS, not an app
What happened — the test, gate by gate

The run, as it happened.

Every number below is from the files on disk, not from a plan.

SCRIPT GATEthree segments, 48 words, on-screen hints kept in their own field.
VOICEcloned professional voice, per segment with context, joined into one 18.2-second WAV; 54 words timed by the transcriber.
STORYBOARD GATEcut into 4 shots at phrase ends (3.5 s · 3.9 s · 4.6 s · 6.3 s); 4 prompts; the validator found no text leaks.
STYLE KEY + IMAGESone key generated from the descriptor; 4 images in parallel with the key as reference — same character, same paper, same lamp.
ASSEMBLYmix at −14 LUFS; one 20-second chunk rendered (543 frames); then re-rendered as three chunks: identical output (SSIM 0.91–0.95, encoder noise only).
MASTER18.10 s video, 18.10 s audio, on-screen words as type. Opened, watched, kept.
contact sheet of the four shots
images/contact-sheet-01.jpg
render QA sheet
renders/dense-sheet.jpg · 1 frame / 2 s
Styles — the machine is style-agnostic

One reference image is the whole style.

A style is three strings and one image: a descriptor (medium, palette, line, protagonist, background, composition), a negative, and a key image generated once from them. Change the file, and the same machine makes a different channel. Six presets ship; a custom style is read from the frames you love.

notebook-stick-figurenotebook-stick-figurethe grammar explainer audiences already reward — paper, ink, stick figure, one warm accent · DEFAULT
flat-vector-editorialflat-vector-editorialfinance / business / tech — navy, cream, coral, no outlines
paper-cutout-collagepaper-cutout-collagehistory / storytelling — cardstock edges, kraft, one mustard accent
chalkboard-lecturechalkboard-lecturescience / how things work — chalk on slate, one pale yellow
soft-watercolor-storysoft-watercolor-storywellness / reflective — washes, sage, rose, ochre
retro-print-halftoneretro-print-halftoneculture / economics — two-ink risograph, halftone
The standard — six laws the line enforces

Why the machine is shaped like this.

01
Script before spend.Every dollar downstream — voice credits, 150 images, a render hour — follows the words. The script gate is not optional.
02
The voice is the clock.Shots are cut from the transcribed audio, never from the script's guess. A voice file you bring enters at station 3.
03
One style key, attached everywhere.Consistency is a reference image, not adjectives. Swap the key and the descriptor and the machine is re-skinned.
04
Text is a layer, never pixels.The validator refuses prompts that ask for words; on-screen text is rendered in a real font — editable, spellable, translatable.
05
Numbers are real or absent.Ideas come from a scraped outlier table; claims carry URLs; the package quotes the CSV. The machine never invents a view count.
06
Generation is disclosed.Every image carries an AI-provenance sidecar; the upload checklist marks synthetic content where YouTube asks.
Your turn

Have something to explain? You have a channel waiting.

One machine, two doors: install it yourself inside the community, or have Nave Partners build and run the channel for your business.

Join the community → Talk to Nave Partners →
STATIONS8 · 3 gates
TEST PIECE18 s · 4 shots · 1 master
RECORDINGS0
FROM IDEA TO FILEone chat, end to end
DIRECTED BY ........ ENRIQUE
PRODUCED BY ........ CLAUDE CODE (the Content OS)
VOICE .............. ELEVENLABS · KOKORO
EARS ............... ASSEMBLYAI
IMAGES ............. NANO BANANA PRO (HIGGSFIELD · KIE.AI)
RENDER ............. HYPERFRAMES
PUBLISH ............ ZERNIO · YOUTUBE STUDIO