bedtime
/ 006 — voice cloning, technically

how it works.

  1. 01
    the reader records 30 seconds of themselves talking.

    minimum 30 seconds, maximum 2 minutes. we provide a phonetic passage and a soft bedtime passage to capture both technical clarity and the register the agent will read in.

  2. 02
    we upload that sample to an audio model.

    ElevenLabs Multilingual v2 is our primary. The voice ID lives in our database. the audio file lives encrypted in S3 with a key Postgres can’t reach.

  3. 03
    the household picks a book.

    public-domain library by default. ~2,000 titles, curated, every one read end-to-end by a human contractor before adding. Family-tier households can also upload PDFs that pass a safety classifier.

  4. 04
    we pre-render the audio, chapter by chapter.

    first chapter ready in about a minute. while it plays we render the rest in the background — we keep a 3-chapter buffer ahead of the listener. subsequent plays of the same book come from cache, instant.

  5. 05
    the player slows down as the story ends.

    a small post-processing pass applies a quieter volume and slower pace across the last 90 seconds. the closing line is the household’s choice. the screen dims to amber, then black.

  6. 06
    every file is watermarked.

    per the AI Voice Provenance Initiative spec. inaudible. traceable to household + reader if the file ever shows up outside the household.

cost. the heaviest line item is TTS — about $0.30 per hour of audio at studio-tier ElevenLabs pricing. a typical session is 7 minutes, so a typical session costs us about 3.5¢ to read.

fallback. a self-hosted stack (XTTS-v2 / F5-TTS on a single GPU) handles overflow from the Family tier and keeps marginal cost ~$0.01 per minute of audio.

voice degradation. every 12 months we email the reader: your voice hasn’t been refreshed in a year. would you like to record again?if they don’t, the clone keeps working. nothing is forced.