Status, 29 September 2026: scripts and optional original non-speech cues are released under the stated licences. No narrated speech is released. Twenty Australian English Polly candidates are stored only in the review-pending folder. A separate local Kokoro test sample was disposable and is not part of this pack.
Rights research for a possible future voice
- Kokoro-82M publisher model card
lists the model as Apache-2.0, explicitly welcomes real-world commercial
deployment, and describes the training audio as permissive or
non-copyrighted. Its associated
voice list
lists
af_heartas American English. The locally downloaded modelkokoro-v1_0.pthhad SHA-256496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4; voice fileaf_heart.pthad SHA-2560ab5709b8ffab19bfd849cd11d98f75b60af7733253ad0d67b12382a102cb4ff. The cache revision wasf3ff3571791e39611d31c381e3a41a3af07b4987. - Output-use boundary: A permissive model-weight licence and permission to deploy the model do not, by themselves, establish who owns each generated waveform or grant us every right needed to relicense one under CC BY 4.0. The model card does not state a separate explicit output-use grant. We therefore make no CC BY claim for an unreleased synthetic narration.
- A Piper lessac voice model card links to Blizzard 2013 training-data terms that permit research use and exclude commercial purposes. That candidate was rejected.
- Amazon Polly states that generated audio
files can be stored and redistributed for any use case. The current
AWS Service Terms, section 50.2
say AI-service output is Your Content, while noting that output may not be
unique. The official voice list
lists Australian English Olivia with neural and generative engines. These
sources provide a clearer route for a future Australian English candidate,
subject to the service terms that apply when it is made and clip-by-clip
human review. After the account session was renewed, twenty Olivia neural
candidate clips were made in
ap-southeast-2from the twenty original scripts. Each file has its own source-text hash, audio hash, engine, voice, region and encoding metadata in the local-onlyreview-pending/polly-olivia-neural/candidate-manifest.json. That directory is Git-ignored and is absent from a clean clone. None has been relicensed or published as a classroom-ready recording.
Quality boundary
The local Kokoro test generated a WAV, but this agent's playback interface did not make audio available to hear. The same constraint applies to the Polly candidate clips. We cannot verify their Australian early-years pronunciation, pacing, clipped words or child-facing tone by listening. No narration should be published until a human listens to every clip against its exact script, checks vocabulary and accent for the intended Australian audience, and confirms a rights basis for the generated output. The scripts may instead be recorded by a consenting original human narrator whose recording rights are obtained for CC BY 4.0 distribution.
Mechanical screening found all twenty candidates to be valid 24 kHz mono
MP3s with peaks below -7 dBFS and no decoded clipping. A local speech
recogniser matched the scripts closely, but this is only a screen for missing
words. One candidate, 06 Container match, transcribed “safe lids” as “safe
leads” and needs particular human pronunciation review. The local-only
review-pending/polly-olivia-neural/REVIEW_QUEUE.md lists all pending clips;
it is Git-ignored and is absent from a clean clone.
The two released sounds are deterministic tones made from numbers by the included code. Their duration, frequency range, sample format, fades and peak level are checked mechanically; human listening is still pending. Neither tone should be placed under speech by default. The user controls device volume and can skip all sounds.
Original commentary © NeuroForgeIO Pty Ltd 2026 · CC BY 4.0 · Note changes if adapted.