PLAY
From zero base, I designed it first, so I created problems and accumulated design debt. But as I hired amazing teammates one by one and built the team together, we started solving those issues. It’s like a chaotic yet heartwarming human drama of a project.
- Year
- 2023 - Present
- Role
- Product Designer
@Supertone AI

Target User
Creators making short videos (up to ~3 min) who generate TTS voiceovers and drop them straight into their edits.


The problem — why TTS was painful
- 1. Onboarding frictionmost tools force sign-up → payment → settings before you hear a single voice.
- 2. Quality vs. controleven natural voices left users without fine control over speed, pitch, and variance.
- 3. Anxiety around cloningVoice cloning needs visibly trustworthy UX for consent, security, and responsibility.
- 4. Multilingual hasslesKeeping one voice's character consistent across languages was confusing.
- 5. Credit/plan fatigueCredits, download limits, commercial use, and attribution were never clear at the right moment.

Core Solutions
- 1) Three‑Step Mental Model: Voice Pick → Type → ListenVoice cards show emotion keywords and example use-cases to cut choice overload.
- 2) Live Emotion/Tone Controls (Pitch Shift / Pitch Variance / Speed)Real-time pitch shift, pitch variance, and speed from a side panel; upload a guide recording and the chosen voice follows it.
- 3) Voice Cloning Onboarding (BETA)Script capture → Quality check → One-time training → Saved to a personal Voice Library, wrapped in an always-visible consent & security layer (encryption, private access, easy deletion).
- 4) Same Character Across Languagesthe language toggle lives inside the same voice-card/script context, with previews for language-specific quirks (numbers, units, emoji).
- 5) In‑Context Guidance on Usage ConditionsCredits, download limits, commercial use, and Free-plan attribution surfaced right in the download modal and header badges.


Key Design Decisions
- Shrink the learning curve with a three-step mental model vs. typical TTS layouts.
- Control-exposure strategy: great defaults for newcomers, instant micro-tuning for power users.
- Ethics as UI — consent, protection, and deletion are always one tap away, not buried in settings.
- Gentle Free-plan attribution reminders at the moment of download, not upfront friction.
- Language switching stays inside the task, never on a separate page.
Impact
The onboarding rework is the change that moved the most: completion rose 19 points, and the shorter path to a first result carried through to how many people came back.
Adoption
- Registered users
- 350,000
- Monthly active users
- 85,000
- 30-day retention
- 24%
- Onboarding completion
- 48% → 67%
- Users creating weekly
- 42%
- How it landedUsers called it "easy and fun," and came back to how quickly they could get to a finished clip — the three-step model doing its job.
- What they asked for nextA wider range of voices and styles, and deeper editing of the generated result.
- What we saw internallyAfter the onboarding rework, time to a first piece of content dropped and drop-off fell with it. Cloning onboarding simplified the reuse loop and improved revisit motivation; holding one character across languages strengthened brand trust.
- Outside the productPicked up as a product example in AI, creator and entertainment press and communities.