PERSO.ai logo

PERSO.ai

Free

Explore Perso AI, the platform for generating AI avatar videos with multilingual support, 1080p export, and chroma key capabilities.

4.5
3
#youtube#twitter
Inputs: text, audio, videoOutputs: video, audio
Type
Saas
Company
PERSO.ai

About PERSO.ai

Perso AI is a 3-in-1 AI audio and video platform combining AI Dubbing, Speech-to-Text, and Audio Separation in a single workflow. It translates and dubs videos into 33+ languages with natural voice cloning and lip dubbing, generates speaker-separated transcripts with automatic speaker diarization in four output formats (XLSX, SRT, VTT, JSON), and isolates individual speaker voices from background audio with dual modes (vocals-only or with reactions preserved). Trusted by 460,000+ users across 80+ countries, powered by the ElevenLabs voice engine (2025 partnership), and ISO/IEC 27001 and KISA ISMS certified. Developed by ESTsoft (est. 1993, KOSDAQ: 047560). Starts at $6.99/month with up to 98% cost savings vs. traditional dubbing studios.

How to Use

Sign up at perso.ai and upload any video or audio file. Choose your workflow — AI Dubbing to translate videos into 33+ languages, Lip Dubbing for natural lip-synced output, Speech-to-Text for speaker-separated transcripts and subtitles (XLSX, SRT, VTT, JSON), or Audio Separation to isolate individual speakers and background audio. Edit the auto-generated script in the real-time editor for instant regeneration, then download or export. Enterprise users can access all capabilities via the API for batch processing.

Key Features

  • AI Dubbing in 33+ languages with voice cloning
  • Lip Dubbing (formerly Lip Sync) — natural mouth-movement alignment
  • Speech-to-Text with automatic speaker diarization (XLSX, SRT, VTT, JSON output)
  • Audio Separation with speaker-level voice isolation
  • Dual background separation modes (vocals-only or with reactions)
  • Custom track combination and export (single merged file)
  • Multi-speaker detection (up to 10 speakers per video)
  • Real-time script editor with instant regeneration
  • Background music and sound effects preservation
  • Enterprise API access + batch video processing
  • ISO/IEC 27001 & KISA ISMS certified data security

Use Cases

  • Corporate L&D Teams — Translate training and onboarding videos. Speech-to-Text auto-generates meeting transcripts with speaker diarization.
  • YouTube Creators — Dub videos into 33+ languages without hiring voice actors. Reach global audiences with localized content.
  • Marketing Agencies — Localize campaign and product-demo videos for international markets while maintaining brand voice.
  • Podcast Producers — Use Audio Separation to extract individual speaker voices, remove background music, or merge selected tracks for post-production.
  • E-Learning Platforms — Translate lecture videos into regional languages. Auto-generate SRT subtitles for accessibility.
  • Media Production — Create dubbed versions of documentary content for international distribution at a fraction of traditional dubbing costs.

Key Features

AI Dubbing in 33+ languages with voice cloning
Lip Dubbing (formerly Lip Sync) — natural mouth-movement alignment
Speech-to-Text with automatic speaker diarization (XLSX, SRT, VTT, JSON output)
Audio Separation with speaker-level voice isolation
Dual background separation modes (vocals-only or with reactions)
Custom track combination and export (single merged file)
Multi-speaker detection (up to 10 speakers per video)
Real-time script editor with instant regeneration
Background music and sound effects preservation
Enterprise API access + batch video processing
ISO/IEC 27001 & KISA ISMS certified data security

Pros & Cons

Pros
  • Comprehensive all-in-one platform for video dubbing and avatars
  • Supports a wide range of languages, facilitating global reach
  • High-quality 1080p export suitable for professional use
  • Chroma key feature allows flexible background manipulation
  • Appears to offer a free tier for initial exploration (limits should be verified)
Cons
  • Free tier likely has usage limits on exports or features
  • Output quality may vary depending on input audio and video quality
  • Requires a stable internet connection for cloud-based processing
  • Detailed pricing for premium plans should be confirmed on the official site
  • May not support all video formats or complex editing workflows

Best For

Corporate L&D Teams — Translate training and onboarding videos. Speech-to-Text auto-generates meeting transcripts with speaker diarization.YouTube Creators — Dub videos into 33+ languages without hiring voice actors. Reach global audiences with localized content.Marketing Agencies — Localize campaign and product-demo videos for international markets while maintaining brand voice.Podcast Producers — Use Audio Separation to extract individual speaker voices, remove background music, or merge selected tracks for post-production.E-Learning Platforms — Translate lecture videos into regional languages. Auto-generate SRT subtitles for accessibility.Media Production — Create dubbed versions of documentary content for international distribution at a fraction of traditional dubbing costs.

Alternatives to PERSO.ai

FAQ

Is PERSO.ai free to use?
Based on available information, PERSO.ai appears to offer a free tier. However, the exact limits on usage, export length, or features should be verified on the official website or current pricing page.
How many languages does PERSO.ai support?
The platform supports over 32 languages for video translation and dubbing. The specific language list should be confirmed in the product documentation.
Can I export videos in 1080p resolution?
Yes, PERSO.ai advertises 1080p export capability. This is likely available on paid plans; free exports may have resolution or watermark restrictions.
Does PERSO.ai provide chroma key (green screen) features?
Yes, chroma key capabilities are included, allowing users to replace or manipulate video backgrounds. The exact implementation should be tested within the platform.
Can I use my own voice for cloning?
PERSO.ai appears to support voice cloning, but the process and allowed usage (e.g., cloning from uploaded audio) should be reviewed in the platform's terms and guides.