Local AI SuiteQwen2.5 + Kokoro + Piper + Auto-Ducking

Scripts, audio, or video into polished content.
100% on your machine.

Generate full videos from raw scripts or voice MP3s, or burn animated captions and secondary overlay badges onto existing MP4s. Complete with Qwen2.5 scene intelligence, Kokoro & Piper voices, and sidechain auto-ducking.

Qwen2.5 Concept Brain64+ Voices & 20 LanguagesSidechain Audio Ducking
SCRIPT-TO-VIDEO9:16 Preview
Qwen2.5
Follow @Creator
Lo-Fi Ambient
-8dB Auto-Duck

Dropascript,Qwen2.5plansthescenes,Kokorovoices,B-rollsynced.

00:00.2 / 00:08.0
Pexels • Pixabay • My Space
WASM Whisper • Sub-millisecond sync • GPU export
Three Flexible Creation Pipelines

Start from a script, an audio recording, or raw video.

JustAddVoice adapts to your creative workflow with three dedicated on-device production pipelines.

Workflow 01

I Have a Script

Paste your text script. The app plans visual scenes, speaks with local studio voices, and stitches auto-paced B-roll.

1
Qwen2.5 Planning: Builds filmable scene prompts.
2
Kokoro/Piper Synthesis: 64+ natural local neural voices.
3
B-Roll Auto-Pacing: Pexels + Pixabay + My Space.
9:16 Shorts / 16:9~15-30s build
Workflow 02

I Have an Audio File

Drop in podcasts, voice memos, or audiobook chapters. Transcribe on-device and turn speech into engaging captioned video.

1
WASM Whisper STT: Sub-millisecond word timestamps.
2
Qwen2.5 Concept Extraction: Filters filler words for clean imagery.
3
Automated Assembly: Paced B-roll sync + subtitle burn.
Up to 15+ min filesNear real-time
Workflow 03

I Have a Video File

Upload raw MP4 videos. Generate animated karaoke captions, edit subtitles live, and attach customizable callout badges.

1
Cinematic Captions: Interactive editor & custom styles.
2
Secondary Overlays: Subscribe badges, callouts, and stickers.
3
GPU Accelerated Export: Burned-in libass subtitles & overlays.
Native GPU CompositingZero Quality Loss
The On-Device Engine Stack

Four specialized local AI engines. One unified desktop app.

No bloated cloud fees. No subscription metering. Just pure on-device artificial intelligence fine-tuned for rapid content production.

Qwen2.5
Local LLM Brain

Qwen2.5 — Scene Planning & Visual Intelligence

The embedded Qwen2.5 model acts as your automated director. It parses your script or audio transcript, extracts key visual concepts, filters out conversational filler words, and constructs precise, filmable stock video queries.

  • Runs on-device via a bundled local server with zero cloud latency
  • Translates abstract narrative ideas into concrete visual queries
  • Optional plug-and-play cloud support (Claude, Grok, OpenAI, Gemini)
Kokoro + Piper
Dual Neural Speech Engines

Kokoro & Piper — 64+ Studio Voices Across 15+ Languages

Generate ultra-natural voiceovers with local Kokoro-82M (24kHz studio quality) and Piper standalone ONNX neural synthesis. Access over 64 distinct voice personalities across English, Spanish, French, German, Italian, and more.

  • 64+ natural voices with adjustable speed, pitch, and cadence
  • Zero per-character API charges or subscription tiers
  • Instant local generation directly to uncompressed audio
FFmpeg Sidechain Ducking
Intelligent Audio Architecture

Ambient Audio Engine & Sidechain Auto-Ducking

Add atmospheric polish with curated offline royalty-free ambient/lo-fi presets ('Lo-Fi Calm', 'Cinematic Ambient', 'Deep Focus') or your own custom audio tracks. Native FFmpeg sidechain compression automatically ducks background music 6–8dB under voiceover.

  • 5 bundled royalty-free CC0 ambient loop presets + custom file upload
  • Real-time Web Audio timeline mixer with adjustable BGM level (12–18%)
  • Native FFmpeg `sidechaincompress` auto-ducking with smooth intro/outro fades
WASM + HY-MT1.5 + Overlays
Subtitles & Canvas Overlays

WASM Whisper, HY-MT1.5 & Secondary Overlays

Extract sub-millisecond word timestamps for karaoke captions, dub across 20 languages with Tencent HY-MT1.5, and drag customizable secondary overlay badges ('Subscribe', 'Follow', custom callouts) directly on the video canvas.

  • WASM Whisper with sub-millisecond word precision
  • Tencent HY-MT1.5 on-device translation for 20 languages
  • Customizable secondary overlay badges composited via GPU FFmpeg
Honest Architecture Matrix

Local AI Core. Optional Hybrid B-Roll.

We don't make false "100% offline" marketing claims. Your entire creation stack — LLM brain, voice models, ambient ducking mixer, translation, and GPU rendering — runs 100% locally on your computer. The internet is only used when you choose to pull fresh stock footage.

Semantic Concept & Visual Query Planning
Alibaba Qwen2.5-3B-Instruct (GGUF / bundled llama-server)
100% Local Device
Runs on your local CPU/GPU. Extracts visual concepts and builds stock footage search prompts without cloud API latency.
Neural Speech Synthesis (TTS)
Kokoro-82M (24kHz) & Piper ONNX (64+ Voices)
100% Local Device
Studio-grade neural voice inference directly on your machine. Zero per-character cloud metering or subscription caps.
Ambient Audio Mixing & Auto-Ducking
Web Audio Mixer + FFmpeg Sidechain Compression
100% Local Device
Offline CC0 ambient music loop presets with native sidechain ducking (-6dB to -8dB under speech) and intro/outro fades.
On-Device Neural Translation & Dubbing
Tencent HY-MT1.5 (1.8B parameter GGUF)
100% Local Device
Translates script sentences into 20 languages with proportional timing alignment and CJK tokenization via Intl.Segmenter.
Speech-to-Text & Subtitle Alignment
Whisper WASM Engine (Sub-millisecond)
100% Local Device
Word-level timestamp extraction in browser sandbox memory. Perfect karaoke highlights for Shorts and Reels.
Video Compositing, Overlays & GPU Export
Native FFmpeg with QuickSync (h264_qsv / libx264)
100% Local Device
Server-native GPU rendering burns libass subtitles and customizable secondary overlay badges directly to high-bitrate MP4.
Personal Media Asset Library ('My Space')
Local Storage & Manifest Indexer
100% Local Device
Your private B-roll clips, custom overlays, and project versions are stored entirely in your local disk workspace.
Stock Footage Discovery (B-Roll)
Pexels & Pixabay API Connectors
Optional Hybrid Cloud
Only accessed when searching for fresh external stock footage. Network dependency shrinks as your local 'My Space' catalog grows.

Complete Audio & Script Privacy

Zero Telemetry Leaks

Your voiceover scripts, synthetic voice models, and audio recordings never leave your machine. No cloud queues, no per-minute compute bills, and no risk of private intellectual property being used for third-party model training.

Growing Local "My Space" Library

Self-Sustaining Catalog

Every B-roll clip fetched from Pexels or Pixabay (or imported from your local footage) is indexed into your local "My Space" library. Future generations automatically reuse matching local assets, making video creation faster over time.

Simple, Transparent Pricing

Start creating with local AI today.

Zero recurring cloud compute charges. Run all local neural voice, translation, and transcription models on your own machine.

Annual Access

Billed annually

€49/ year

Flexible yearly access with continuous updates.

What's included:

  • Full desktop app for Windows
  • All local AI engines (Qwen2.5, Kokoro, Piper, Whisper)
  • Continuous engine and feature updates
  • Pexels & Pixabay stock B-roll integration
  • Standard customer support
Get Annual Access

Cancel anytime. Unmetered on-device generations.

Best Value

Lifetime License

One-time payment

€149one-time

Pay once, own the full software suite forever.

Launch Special: First 100 Lifetime licenses are €99
Use codeat checkout.

Everything in Annual, plus:

  • Own it forever with zero recurring subscriptions
  • All future updates, new model integrations & tools
  • Priority customer support & direct roadmap feedback
  • Upcoming ambient music engine & sidechain ducking
Get Lifetime License

Instant license key delivery • 100% offline engine

Note for Windows Users:

Because JustAddVoice is a brand-new release, Windows SmartScreen may flag the initial download. To install safely, simply click "More Info""Run Anyway".