Scripts, audio, or video into polished content.
100% on your machine.
Generate full videos from raw scripts or voice MP3s, or burn animated captions and secondary overlay badges onto existing MP4s. Complete with Qwen2.5 scene intelligence, Kokoro & Piper voices, and sidechain auto-ducking.
Dropascript,Qwen2.5plansthescenes,Kokorovoices,B-rollsynced.
Start from a script, an audio recording, or raw video.
JustAddVoice adapts to your creative workflow with three dedicated on-device production pipelines.
I Have a Script
Paste your text script. The app plans visual scenes, speaks with local studio voices, and stitches auto-paced B-roll.
I Have an Audio File
Drop in podcasts, voice memos, or audiobook chapters. Transcribe on-device and turn speech into engaging captioned video.
I Have a Video File
Upload raw MP4 videos. Generate animated karaoke captions, edit subtitles live, and attach customizable callout badges.
Four specialized local AI engines. One unified desktop app.
No bloated cloud fees. No subscription metering. Just pure on-device artificial intelligence fine-tuned for rapid content production.
Qwen2.5 — Scene Planning & Visual Intelligence
The embedded Qwen2.5 model acts as your automated director. It parses your script or audio transcript, extracts key visual concepts, filters out conversational filler words, and constructs precise, filmable stock video queries.
- Runs on-device via a bundled local server with zero cloud latency
- Translates abstract narrative ideas into concrete visual queries
- Optional plug-and-play cloud support (Claude, Grok, OpenAI, Gemini)
Kokoro & Piper — 64+ Studio Voices Across 15+ Languages
Generate ultra-natural voiceovers with local Kokoro-82M (24kHz studio quality) and Piper standalone ONNX neural synthesis. Access over 64 distinct voice personalities across English, Spanish, French, German, Italian, and more.
- 64+ natural voices with adjustable speed, pitch, and cadence
- Zero per-character API charges or subscription tiers
- Instant local generation directly to uncompressed audio
Ambient Audio Engine & Sidechain Auto-Ducking
Add atmospheric polish with curated offline royalty-free ambient/lo-fi presets ('Lo-Fi Calm', 'Cinematic Ambient', 'Deep Focus') or your own custom audio tracks. Native FFmpeg sidechain compression automatically ducks background music 6–8dB under voiceover.
- 5 bundled royalty-free CC0 ambient loop presets + custom file upload
- Real-time Web Audio timeline mixer with adjustable BGM level (12–18%)
- Native FFmpeg `sidechaincompress` auto-ducking with smooth intro/outro fades
WASM Whisper, HY-MT1.5 & Secondary Overlays
Extract sub-millisecond word timestamps for karaoke captions, dub across 20 languages with Tencent HY-MT1.5, and drag customizable secondary overlay badges ('Subscribe', 'Follow', custom callouts) directly on the video canvas.
- WASM Whisper with sub-millisecond word precision
- Tencent HY-MT1.5 on-device translation for 20 languages
- Customizable secondary overlay badges composited via GPU FFmpeg
Local AI Core. Optional Hybrid B-Roll.
We don't make false "100% offline" marketing claims. Your entire creation stack — LLM brain, voice models, ambient ducking mixer, translation, and GPU rendering — runs 100% locally on your computer. The internet is only used when you choose to pull fresh stock footage.
Complete Audio & Script Privacy
Zero Telemetry LeaksYour voiceover scripts, synthetic voice models, and audio recordings never leave your machine. No cloud queues, no per-minute compute bills, and no risk of private intellectual property being used for third-party model training.
Growing Local "My Space" Library
Self-Sustaining CatalogEvery B-roll clip fetched from Pexels or Pixabay (or imported from your local footage) is indexed into your local "My Space" library. Future generations automatically reuse matching local assets, making video creation faster over time.
Start creating with local AI today.
Zero recurring cloud compute charges. Run all local neural voice, translation, and transcription models on your own machine.
Annual Access
Billed annually
Flexible yearly access with continuous updates.
What's included:
- Full desktop app for Windows
- All local AI engines (Qwen2.5, Kokoro, Piper, Whisper)
- Continuous engine and feature updates
- Pexels & Pixabay stock B-roll integration
- Standard customer support
Cancel anytime. Unmetered on-device generations.
Lifetime License
One-time payment
Pay once, own the full software suite forever.
Everything in Annual, plus:
- Own it forever with zero recurring subscriptions
- All future updates, new model integrations & tools
- Priority customer support & direct roadmap feedback
- Upcoming ambient music engine & sidechain ducking
Instant license key delivery • 100% offline engine
Note for Windows Users:
Because JustAddVoice is a brand-new release, Windows SmartScreen may flag the initial download. To install safely, simply click "More Info" → "Run Anyway".