Skip to content
Max Sungmin Kim

PLAY

From zero base, I designed it first, so I created problems and accumulated design debt. But as I hired amazing teammates one by one and built the team together, we started solving those issues. It’s like a chaotic yet heartwarming human drama of a project.

Year
2023 - Present
Role
Product Designer
@Supertone AI
PLAY product screen

Target User

Creators making short videos (up to ~3 min) who generate TTS voiceovers and drop them straight into their edits.

PLAY MVP version, 2024
MVP Version (2024)
PLAY service version, 2025
Service Ver. 2025

The problem — why TTS was painful

  • 1. Onboarding frictionmost tools force sign-up → payment → settings before you hear a single voice.
  • 2. Quality vs. controleven natural voices left users without fine control over speed, pitch, and variance.
  • 3. Anxiety around cloningVoice cloning needs visibly trustworthy UX for consent, security, and responsibility.
  • 4. Multilingual hasslesKeeping one voice's character consistent across languages was confusing.
  • 5. Credit/plan fatigueCredits, download limits, commercial use, and attribution were never clear at the right moment.
Credit and plan surfaces

Core Solutions

  • 1) Three‑Step Mental Model: Voice Pick → Type → ListenVoice cards show emotion keywords and example use-cases to cut choice overload.
  • 2) Live Emotion/Tone Controls (Pitch Shift / Pitch Variance / Speed)Real-time pitch shift, pitch variance, and speed from a side panel; upload a guide recording and the chosen voice follows it.
  • 3) Voice Cloning Onboarding (BETA)Script capture → Quality check → One-time training → Saved to a personal Voice Library, wrapped in an always-visible consent & security layer (encryption, private access, easy deletion).
  • 4) Same Character Across Languagesthe language toggle lives inside the same voice-card/script context, with previews for language-specific quirks (numbers, units, emoji).
  • 5) In‑Context Guidance on Usage ConditionsCredits, download limits, commercial use, and Free-plan attribution surfaced right in the download modal and header badges.
Voice controls panel
Voice cloning onboarding

Key Design Decisions

  • Shrink the learning curve with a three-step mental model vs. typical TTS layouts.
  • Control-exposure strategy: great defaults for newcomers, instant micro-tuning for power users.
  • Ethics as UI — consent, protection, and deletion are always one tap away, not buried in settings.
  • Gentle Free-plan attribution reminders at the moment of download, not upfront friction.
  • Language switching stays inside the task, never on a separate page.

Impact

The onboarding rework is the change that moved the most: completion rose 19 points, and the shorter path to a first result carried through to how many people came back.

Adoption

Registered users
350,000
Monthly active users
85,000
30-day retention
24%
Onboarding completion
48% → 67%
Users creating weekly
42%
  • How it landedUsers called it "easy and fun," and came back to how quickly they could get to a finished clip — the three-step model doing its job.
  • What they asked for nextA wider range of voices and styles, and deeper editing of the generated result.
  • What we saw internallyAfter the onboarding rework, time to a first piece of content dropped and drop-off fell with it. Cloning onboarding simplified the reuse loop and improved revisit motivation; holding one character across languages strengthened brand trust.
  • Outside the productPicked up as a product example in AI, creator and entertainment press and communities.