Consent and provenance in voice cloning
A few seconds of reference audio is now enough to reproduce a speaking voice. The interesting questions are all about record-keeping.
Reference-based voice synthesis works from a short sample: a speaker reads a few seconds of audio and the system produces new speech in that voice. The technical barrier that once made this a specialist activity is gone, and what remains is a set of record-keeping problems that are unglamorous and entirely tractable.
The framing that helps most is to stop asking whether cloning a voice is permitted in the abstract and start asking two narrower questions. Can you demonstrate, later, that the speaker agreed to this use? And can a listener, or you, later determine that a given recording was synthesised?
What a consent record has to contain
A consent record that is worth keeping identifies the speaker, the date, the specific recording used as reference, and the scope of use that was agreed. Scope is the part most often left vague and the part most often disputed: a speaker who agrees to a voice being used for one project has not thereby agreed to it being used for another, and a record that does not state the boundary cannot be relied on to establish it.
Duration matters as much as scope. A consent that never expires is a consent that cannot be withdrawn, which in practice means the recording has to be treated as permanent. Recording a review date, and a route by which a speaker can ask for the reference audio and derived material to be withdrawn, is the difference between a policy and a gesture.
Permission that was never written down is indistinguishable, later, from permission that was never given.
Provenance travels badly
Metadata attached to an audio file is stripped by nearly every platform that re-encodes it. Any provenance scheme that depends only on a file header will survive the first upload and no further. This is why marking that is carried in the audio signal itself, and cryptographic manifests kept alongside the distribution, are usually discussed together rather than as alternatives.
Neither is a complete answer. Signal marking survives re-encoding and can be degraded deliberately; manifests are robust but only if the verifying party bothers to check. Practically, they cover different attacks, and a project that keeps both is in a defensible position.
Internal traceability
Independent of what ships publicly, keeping a link from every generated file back to the reference audio and the consent record that authorised it is cheap at build time and impossible to reconstruct afterwards. When a question arises about a specific clip, the useful answer is a record, not a recollection.
The general principle is unremarkable and worth stating plainly: the controls that work are the ones written down before they are needed.