# Narration production notes

**Status, 29 September 2026:** scripts and optional original non-speech cues
are released under the stated licences. No narrated speech is released.
Twenty Australian English Polly candidates are stored only in the
[review-pending folder](review-pending/README.md). A separate local Kokoro
test sample was disposable and is not part of this pack.

## Rights research for a possible future voice

- [Kokoro-82M publisher model card](https://huggingface.co/hexgrad/Kokoro-82M)
  lists the model as Apache-2.0, explicitly welcomes real-world commercial
  deployment, and describes the training audio as permissive or
  non-copyrighted. Its associated
  [voice list](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md)
  lists `af_heart` as American English. The locally downloaded model
  `kokoro-v1_0.pth` had SHA-256
  `496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4`;
  [voice file](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/voices/af_heart.pt)
  `af_heart.pt` had SHA-256
  `0ab5709b8ffab19bfd849cd11d98f75b60af7733253ad0d67b12382a102cb4ff`.
  The cache revision was `f3ff3571791e39611d31c381e3a41a3af07b4987`.
- **Output-use boundary:** A permissive model-weight licence and permission to
  deploy the model do not, by themselves, establish who owns each generated
  waveform or grant us every right needed to relicense one under CC BY 4.0.
  The model card does not state a separate explicit output-use grant. We
  therefore make no CC BY claim for an unreleased synthetic narration.
- A [Piper lessac voice model card](https://huggingface.co/rhasspy/piper-voices/blob/main/en/en_US/lessac/medium/MODEL_CARD)
  links to [Blizzard 2013 training-data terms](https://www.cstr.ed.ac.uk/projects/blizzard/2013/lessac_blizzard2013/license.html)
  that permit research use and exclude commercial purposes. That candidate
  was rejected.
- [Amazon Polly](https://aws.amazon.com/polly/) states that generated audio
  files can be stored and redistributed for any use case. The current
  [AWS Service Terms, section 50.2](https://aws.amazon.com/service-terms/)
  say AI-service output is Your Content, while noting that output may not be
  unique. The [official voice list](https://docs.aws.amazon.com/polly/latest/dg/available-voices.html)
  lists Australian English Olivia with neural and generative engines. These
  sources provide a clearer route for a future Australian English candidate,
  subject to the service terms that apply when it is made and clip-by-clip
  human review. After the account session was renewed, twenty Olivia neural
  candidate clips were made in `ap-southeast-2` from the twenty original
  scripts. Each file has its own source-text hash, audio hash, engine, voice,
  region and encoding metadata in the local-only
  `review-pending/polly-olivia-neural/candidate-manifest.json`. That directory
  is Git-ignored and is absent from a clean clone. None has been relicensed
  or published as a classroom-ready recording.

## Quality boundary

The local Kokoro test generated a WAV, but this agent's playback interface
did not make audio available to hear. The same constraint applies to the Polly
candidate clips. We cannot verify their Australian early-years pronunciation,
pacing, clipped words or child-facing tone by listening. No narration should
be published until a human listens to **every**
clip against its exact script, checks vocabulary and accent for the intended
Australian audience, and confirms a rights basis for the generated output.
The scripts may instead be recorded by a consenting original human narrator
whose recording rights are obtained for CC BY 4.0 distribution.

Mechanical screening found all twenty candidates to be valid 24 kHz mono
MP3s with peaks below -7 dBFS and no decoded clipping. A local speech
recogniser matched the scripts closely, but this is only a screen for missing
words. One candidate, 06 Container match, transcribed “safe lids” as “safe
leads” and needs particular human pronunciation review. The local-only
`review-pending/polly-olivia-neural/REVIEW_QUEUE.md` lists all pending clips;
it is Git-ignored and is absent from a clean clone.

The two released sounds are deterministic tones made from numbers by the
included code. Their duration, frequency range, sample format, fades and peak
level are checked mechanically; human listening is still pending. Neither
tone should be placed under speech by default. The user controls device
volume and can skip all sounds.

Original commentary © NeuroForgeIO Pty Ltd 2026 ·
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) · Note changes if adapted.
