Problem
How can the site replace elementary one-liners and nominal duration labels with genuine C1+ long-form material and dependable Japanese and English playback?
Role
Lesson design, source policy, scheduled generation, speech playback, archive UI, and implementation
Process
- Use one complete C1+ Japanese speech and C1+ English speech plus small talk, with tested minimum character and word counts
- Lazy-load Piper Plus only when playback is requested, synthesize natural Japanese and English locally, cache the model, and reserve system speech as a last fallback
- Archive lessons by Tokyo date in Durable Object v2, isolate obsolete short-form records, and keep past sessions selectable
Outcome
The working version enforces at least 4,000 Japanese characters and 3,000 English words, with natural neural speech, speed controls, paragraph playback, and dated archives.
Reflection
Practice time must be supported by real content volume and a protocol, not a label; speech playback must also account for asynchronous voices and browser length limits.