how it works.
- 01the reader records 30 seconds of themselves talking.
minimum 30 seconds, maximum 2 minutes. we provide a phonetic passage and a soft bedtime passage to capture both technical clarity and the register the agent will read in.
- 02we upload that sample to an audio model.
ElevenLabs Multilingual v2 is our primary. The voice ID lives in our database. the audio file lives encrypted in S3 with a key Postgres can’t reach.
- 03the household picks a book.
public-domain library by default. ~2,000 titles, curated, every one read end-to-end by a human contractor before adding. Family-tier households can also upload PDFs that pass a safety classifier.
- 04we pre-render the audio, chapter by chapter.
first chapter ready in about a minute. while it plays we render the rest in the background — we keep a 3-chapter buffer ahead of the listener. subsequent plays of the same book come from cache, instant.
- 05the player slows down as the story ends.
a small post-processing pass applies a quieter volume and slower pace across the last 90 seconds. the closing line is the household’s choice. the screen dims to amber, then black.
- 06every file is watermarked.
per the AI Voice Provenance Initiative spec. inaudible. traceable to household + reader if the file ever shows up outside the household.
cost. the heaviest line item is TTS — about $0.30 per hour of audio at studio-tier ElevenLabs pricing. a typical session is 7 minutes, so a typical session costs us about 3.5¢ to read.
fallback. a self-hosted stack (XTTS-v2 / F5-TTS on a single GPU) handles overflow from the Family tier and keeps marginal cost ~$0.01 per minute of audio.
voice degradation. every 12 months we email the reader: your voice hasn’t been refreshed in a year. would you like to record again?if they don’t, the clone keeps working. nothing is forced.