top of page

1. Onboarding friction
most tools force sign-up → payment → settings before you hear a single voice.
2. Quality vs. control
even natural voices left users without fine control over speed, pitch, and variance.
3. Anxiety around cloning
Voice cloning needs visibly trustworthy UX for consent, security, and responsibility.
4. Multilingual hassles
Keeping one voice's character consistent across languages was confusing.
5. Credit/plan fatigue
Credits, download limits, commercial use, and attribution were never clear at the right moment.
The problem — why TTS was painful
-
Shrink the learning curve with a three-step mental model vs. typical TTS layouts.
-
Control-exposure strategy: great defaults for newcomers, instant micro-tuning for power users.
-
Ethics as UI — consent, protection, and deletion are always one tap away, not buried in settings.
-
Gentle Free-plan attribution reminders at the moment of download, not upfront friction.
-
Language switching stays inside the task, never on a separate page.
Key Design Decisions
Target User
Creators making short videos (up to ~3 min) who generate TTS voiceovers and drop them straight into their edits.



MVP Version (2024)

Service Ver. 2025

Impact
-
Path to first synthesis became clearly faster in usability tests; users called the emotion/speed controls "easy and fun."
-
Cloning onboarding simplified the reuse loop and improved revisit motivation.
-
Consistent character across languages strengthened brand trust.
-
Tracked via: time-to-first-value (TTFV), first-download conversion, cloning-onboarding completion, multilingual-toggle usage, free→paid conversion, and support-ticket volume.
PLAY
From zero base, I designed it first, so I created problems and accumulated design debt. But as I hired amazing teammates one by one and built the team together, we started solving those issues. It’s like a chaotic yet heartwarming human drama of a project.
Year
2023 - Present
Role
Product Designer
@Supertone AI
bottom of page