Content Creator Studio · Self-hosted

Turn the PC you already own into a video studio that never sleeps.

Feed it images and a script. It writes the captions, narrates them in a real voice, adds motion, and renders a ready-to-post MP4 — on your own GPU, with no per-video fee and nothing uploaded to anyone's cloud.

Self-hosted via Docker Compose on your own NVIDIA GPU — not a cloud render queue.

Your GPU is idle while your queue keeps growing

What's slowing you down

One video is 20 minutes of the same clicks

Cut, caption, position the text, export — the same manual pass on every single upload.

Captions retyped by hand

Watching back the whole clip just to transcribe what was already said out loud.

Cloud render fees that scale with your ambition

The more you post, the more a per-minute or per-render bill grows — right when volume should be getting cheaper, not more expensive.

More output, without more hours or a bigger bill

What you get

01

Volume without extra hours

A render is a job, not a sit-down session — queue it, watch it progress, come back to a finished file.

How: background jobs stream live progress over the same connection the dashboard uses to poll status.
02

Zero marginal cost per video

Encoding, transcription and narration all run on hardware you already own — the only costs are the ones you choose to opt into.

How: NVENC video encode, CUDA Whisper transcription, and three local text-to-speech engines — no metered API in the critical path.
03

Production that runs unattended

Set a schedule once; it keeps producing while you're doing anything else.

How: a project automation loop ticks every minute and advances your project one step per run — generate, then export.

From raw material to finished MP4

How it works

1

Bring your images & script

Upload stills or clips, or pull stock images — then write or paste the narration script.

2

Pick a voice

Three local text-to-speech engines, including Vietnamese and voice cloning — all rendered on your own GPU.

3

Choose motion, ratio & captions

Ken Burns direction, aspect ratio, caption style and position — with a live preview before you commit.

4

Render & collect the file

Watch it progress stage by stage, then pull the MP4 and its SRT sidecar from your media library.

What the studio actually looks like

See it

Representative mockups of the real dashboard — not stock photos.

Tools

Video GeneratorImages + narration → captioned MP4
Text to SpeechKokoro, VieNeu, Chatterbox — local
Subtitle VideoWhisper captions, editable, burned in
Sound WaveformAudio-reactive visualizer video

Seven tool cards, each an independent, focused workflow.

Caption style editor

FontBe Vietnam Pro
Size42px
PositionBottom center
In / out effectFade
the sunrise hit different today

The preview renders through the same styling engine as the export.

Waveform visualizer

4 modes × 10 palettes, generated as a real second video layer from the audio.

Export logs

ep12-morning-routine.mp41080×1920 · 41s · rendered in 22.4s
done
ep13-city-walk.mp41080×1920 · queued
processing

Every render keeps a per-stage timing breakdown, not just a pass/fail.

The pipeline, in full

Everything inside

Images + voiceover → captioned MP4

The whole pipeline in one pass — no manual timeline editing required.

Ken Burns motion, 6 directions + fit-screen

4× supersampled so slow zooms move smoothly instead of juddering.

6 aspect ratios, up to 4K

9:16, 16:9, 1:1, 4:5, 4:3, 21:9 — four quality tiers each.

NVENC encode, automatic CPU fallback

Uses your GPU when it can, degrades gracefully when it can't.

Editable Whisper captions

Review and fix the transcript before it's ever burned in — line by line, with an audio scrub.

Caption translation + dual subtitles

EN, VI, JA, ZH, KO — original and translated tracks burned at two positions at once.

3 local voice engines

Kokoro (multi-language), VieNeu (Vietnamese, with cloning), Chatterbox (cloning) — all self-hosted.

Audio-reactive waveform videos

4 modes × 10 palettes, over an image, video or solid background.

Intro cards, end CTAs, timers & overlays

Countdown or count-up timers, multiple free-standing text messages, per-clip fades.

Ear-safe music mixing

Background music ducked under narration and normalized to −14 LUFS / −1 dBTP.

ComfyUI workflow import

Bring your own image/video workflows and run them from inside the studio.

Scheduled production

A project automation loop that generates and exports on the schedule you set.

Honest about the setup — it's self-hosted

What it takes to run

Docker Compose

Runs as a set of containers on your machine or a home server — docker compose up and it's live.

One NVIDIA GPU

Powers NVENC encoding, CUDA transcription and every local voice — recent cards (incl. RTX 50-series) are supported.

Also runs CPU-only

Slower, but nothing is hard-blocked on a GPU — every encode and transcription step has a CPU fallback.

Licensed per install

A license key verified against OurAI Tech activates the product — issued when you sign up.

There is no one-click installer yet — set up is Docker Compose on infrastructure you control, not a downloadable desktop app.

Straight answers

FAQ

Do I need a GPU?

Not strictly — everything falls back to CPU. But a GPU is what makes encoding, transcription and local voices fast, and it's the whole point of the "your own hardware, no per-render fee" pitch.

What does a video cost to make?

Rendering, encoding, transcription and all three built-in voices run on your own GPU — no per-video fee. The only metered costs are optional hosted AI providers (for AI image/video generation or caption translation) that you opt into with your own API key.

Does my footage leave my machine?

Not for the core pipeline — rendering, transcription and local voice all run in your own containers. It only leaves your machine if you deliberately turn on an optional hosted provider.

Can it publish to TikTok or YouTube for me?

No — it produces the finished MP4 and keeps a publish log for your records, but uploading to a platform is still a manual step today.

Does it do Vietnamese voice and captions?

Yes — a dedicated Vietnamese text-to-speech engine with voice cloning, and captions with verified Vietnamese font coverage and translation to and from Vietnamese.

Can I use my own AI image models?

Yes — import a ComfyUI workflow (API format) and run it from inside the studio, alongside optional OpenAI and Google image/video providers.

Ready when you are

Put your GPU to work while you do anything else

Set it up once on your own hardware — every render after that costs you nothing but electricity.

Back to Home