Performed using NEO AI Engineer
If you ship TTS or voice cloning, you eventually need a straight answer: which model sounds natural, stays intelligible, actually clones the reference speaker, and still runs at a usable speed on CPU. This post is that answer first, then the methodology and the three bugs that made earlier publishes untrustworthy.
Lead result: Audio8 now clones per speaker. Identity is a near-tie with XTTS: ECAPA 0.5454 vs 0.5569 (10/22 vs 12/22 speakers; natural ceiling 0.6964). Matched >> mismatch on sent1 (0.4885 vs 0.1130, gap +0.3755). Among cloners on this CPU, Audio8 is the practical pick (UTMOS 3.3657, 19/22; honest RTF 1.4214 vs XTTS 5.3112). The Audio8 re-synth was registration-only — same 264-row VCTK manifest, same references, same test texts; only the wrapper now registers voice_name=speaker_id. XTTS / Kokoro / Pocket TTS were not re-run. The stuck-voice row (SIM 0.1580) is a short postmortem only.
Results
All models + natural ceiling (264 unique sentences each)
From output/aggregated_results.csv:
| Model | Speaker sim ↑ | UTMOS ↑ | WER ↓ | CER ↓ | RTF ↓ | Wall (s) |
|---|---|---|---|---|---|---|
| Pocket TTS | 0.1402 | 3.4138 | 0.0057 | 0.0020 | 0.2433 | 1.0047 |
| Kokoro-82M | 0.0427 | 3.6899 | 0.0051 | 0.0016 | 0.1238 | 0.5121 |
| Audio8-0.6b | 0.5454 | 3.3657 | 0.0063 | 0.0018 | 1.4214 | 5.7269 |
| XTTS-v2 | 0.5569 | 3.2223 | 0.0091 | 0.0029 | 5.3112 | 21.2623 |
| VCTK natural | 0.6964 | 3.5823 | 0.0355 | 0.0221 | — | — |
Audio8 RTF 1.4214 amortizes 22 registrations over 264 utterances (sent1 2.7441; remaining 242 clips 1.3011). The stuck-voice row (0.1580 / 1.3953) is a postmortem only.
Cross-session latency: Audio8 wall / RTF is from the 2026-08-15 registration-only re-synth session. XTTS / Kokoro / Pocket RTF are from the original 2026-08-06 sweep. Same CPU-only box (8 cores, ~62 GB RAM, no GPU), not a same-process A/B. Directionally Audio8 is still ~3.7× faster than XTTS on this machine.
Four-model comparison on UTMOS naturalness, WER intelligibility, and RTF latency
Figure: Aggregate UTMOS, WER, and RTF across 264 unique sentences per model. Dashed line is VCTK-natural UTMOS (3.58). Kokoro leads naturalness and speed. WER is demoted (Whisper on short lines). RTF = wall / duration.
CPU wall-clock latency per utterance for Pocket TTS, Kokoro, Audio8, and XTTS
Figure: Mean wall-clock seconds per utterance on the CPU-only evaluation box.
All 1056 synth rows passed the energy gate with non-null metrics. Natural RTF / wall are blank (no synthesis).
Four-model comparison (Audio8 is a real cloner)
Read the table as two cloners plus two fixed-voice controls, not four cloners.
| Role | Models | What the CSVs say |
|---|---|---|
| Cloners | Audio8, XTTS-v2 | Identity near-tie: Audio8 0.5454 vs XTTS 0.5569 (10/22 vs 12/22; natural ceiling 0.6964). Among cloners Audio8 is the CPU pick: UTMOS 3.3657 (19/22) and honest RTF 1.4214 vs XTTS 5.3112. |
| Fixed-voice controls | Pocket TTS (anna), Kokoro (af_heart) | SIM 0.1402 / 0.0427 sits at the mismatch floor. Kokoro still wins catalog naturalness (3.6899) and speed (RTF 0.1238). |
Do not rank Pocket or Kokoro on identity. Do not treat the stuck-voice Audio8 row (SIM 0.1580) as a current score.
Intended cloners: Audio8 vs XTTS (528 rows)
| Metric | Audio8 | XTTS | Δ (A8 − XTTS) | Speaker-level 95% CI | Takeaway |
|---|---|---|---|---|---|
| Speaker sim | 0.5454 | 0.5569 | −0.0116 | [−0.0336, +0.0130] includes 0 | Near-tie. Audio8 10/22, XTTS 12/22 |
| UTMOS | 3.3657 | 3.2223 | +0.1434 | [+0.0619, +0.2228] excludes 0 | Audio8 19/22 (22 clones vs 22 clones) |
| WER | 0.0063 | 0.0091 | −0.0028 | [−0.0072, +0.0016] includes 0 | Demoted. Does not separate models |
| RTF | 1.4214 | 5.3112 | −3.8898 | — | Honest Audio8 RTF (22 regs amortized) |
Paired bootstrap is speaker-level (resample 22 speakers, B=10000, seed 42). Source: output/stats_bootstrap.json.
p226 UTMOS: p226 is the speaker Audio8 was stuck on in the first publish (every later “clone” still sounded like p226). After the registration-only re-synth, p226 UTMOS is Audio8 3.170 vs XTTS 3.148 (+0.022) — a small win, not the source of the overall +0.1434. The naturalness lead is now 19/22 speakers on 22 clones vs 22 clones, not a fixed p226-like voice scoring well against XTTS clones.
Intended-cloner comparison: Audio8 vs XTTS on UTMOS, WER, RTF, and speaker similarity
Figure: Head-to-head on the two intended cloners (528 rows). Identity is a near-tie. WER is not a ranking. UTMOS favors Audio8 among cloners.
Per-speaker UTMOS naturalness for Audio8 vs XTTS across 22 VCTK speakers
Figure: Per-speaker mean UTMOS. Audio8 wins 19/22 speakers after the re-synth (22 clones vs 22 clones).
Per-speaker WER intelligibility for Audio8 vs XTTS across 22 VCTK speakers
Figure: Per-speaker mean WER (%). Absolute rates are small on short sentences. Do not rank cloners from this column.
Per-speaker ECAPA speaker similarity for Audio8 vs XTTS across 22 VCTK speakers
Figure: Per-speaker mean ECAPA cosine. Audio8 wins 10/22, XTTS 12/22. Dashed line is the VCTK-natural mean (0.6964).
How to read the columns
Speaker similarity
This column now ranks. Natural same-speaker speech sits at 0.6964. XTTS is 0.5569 (n=264). Audio8 is 0.5454 — a near-tie (10/22 vs 12/22). The SIM Δ CI includes 0, so do not call a winner on identity. XTTS 0.4896 / Audio8 0.4885 are the n=22 sent1 matched means. Kokoro is 0.0427 overall.
Speaker similarity (ECAPA-TDNN) for four models plus the VCTK-natural ceiling
Figure: ECAPA-TDNN cosine. Both cloners sit in the identity band under the natural ceiling. Kokoro is near chance, as a fixed female voice should be.
Naturalness
Kokoro is the clear leader (UTMOS 3.6899) if you can accept a fixed voice. It also outscores the 264-clip natural mean (3.5823) — a predictor quirk, called out rather than hidden. Among cloners, Audio8 leads (3.3657 vs 3.2223; 19/22 speakers). The UTMOS Δ CI excludes 0.
Intelligibility (demoted)
WER does not separate these models. Natural 0.0355 vs synth 0.0051–0.0091 is Whisper small rewarding hyper-articulation on short VCTK lines. The WER Δ CI includes 0. Do not lead with WER.
Latency on CPU
Kokoro is comfortably real-time (RTF 0.1238). Pocket TTS is also real-time (0.2433). Audio8 is about 1.4× real-time as a working cloner (honest RTF 1.4214, 22 regs amortized). XTTS is about 5.3× real-time (~21 s wall per utterance on this machine). If your product is interactive CPU TTS and you need XTTS-level identity, budget for a GPU path or prefer Audio8 on this box.
Recommendations
| Goal | Prefer | Reason (CSV) |
|---|---|---|
| Best sounding fixed voice + speed on CPU | Kokoro-82M | Highest UTMOS (3.6899) and lowest RTF (0.1238) |
| Zero-shot identity (both cloners work) | Audio8 or XTTS | Near-tie: 0.5454 vs 0.5569; 10/22 vs 12/22; SIM CI includes 0; natural 0.6964 |
| Zero-shot + CPU latency among cloners | Audio8-0.6b | Honest RTF 1.4214 vs XTTS 5.3112; UTMOS 3.3657 (19/22; CI excludes 0) |
| Fast catalog TTS when identity is not required | Kokoro or Pocket TTS | Real-time RTF; SIM near the mismatch floor |
Caveats worth keeping in the open: objective-only (no human MOS), Whisper small for WER, Pocket/Kokoro not cloners, stuck-voice Audio8 is a postmortem only (wrapper is fixed; published row is the registration-only re-synth), XTTS peak memory not measured in-process, speaker-level paired bootstrap is in output/stats_bootstrap.json (UTMOS CI excludes 0; SIM and WER CIs include 0), ECAPA is one VoxCeleb-trained embedding. Figures are regenerated from the post-fix CSVs.
Full tables and methodology live in output/report.md in the repo.
Methodology
Why this evaluation exists
Vendor demos and single-speaker samples are easy to game. A fair comparison needs:
- Many speakers, not one celebrity voice
- Gender balance and accent diversity
- Held-out text that never appears in the reference clip
- Metrics that separate cloning fidelity, naturalness, intelligibility, and latency
- An honest split between true zero-shot cloners and fixed-voice TTS
This case study was produced end to end by Neo, an autonomous AI engineering agent you can run from VS Code or Cursor.
Models under test
| Model | Role in this study | Zero-shot from reference audio? |
|---|---|---|
| Pocket TTS | Catalog / fallback path | Cloning weights are HF-gated; run used fixed voice anna |
| Kokoro-82M | Fast fixed-voice baseline | No. Fixed embeddings only (af_heart) |
| Audio8-TTS-0.6b | ONNX zero-shot cloner | Yes after wrapper fix (voice_name=speaker_id) |
| XTTS-v2 | Coqui zero-shot cloner | Yes (speaker_wav) |
| VCTK natural | Same-speaker ceiling | Not a model. Held-out clip vs the speaker’s reference |
Audio8 and XTTS are the two intended cloners; both now clone. Pocket TTS and Kokoro are fixed-voice baselines and negative controls for speaker similarity. If a SIM metric says Kokoro matches eleven male VCTK speakers at 0.98, the metric is wrong.
Metrics
All scoring is automatic. No MOS listening panel in this run.
- Speaker similarity (published) — ECAPA-TDNN embeddings from
speechbrain/spkrec-ecapa-voxceleb, cosine similarity, both streams at 16 kHz. Replaced an invalid WavLM-large x-vector column. - Naturalness (UTMOSv2) — Predicted MOS on a 1–5 scale. Fixed-voice Kokoro leads this column. On the full 264 natural clips the mean is 3.5823 — below Kokoro 3.6899 — so do not treat natural UTMOS as a hard ceiling.
- Intelligibility (WER / CER) — faster-whisper
small(CTranslate2, CPU int8) → normalized text → jiwer. Natural VCTK WER is 0.0355 under this ASR; that is Whisper-on-real-VCTK, not a model ranking fail. - Latency — Wall-clock seconds and RTF = wall / audio duration (not the reciprocal). Peak memory for XTTS is not trustworthy in the parent process because XTTS runs in an isolated subprocess (
venv_xtts).
Dataset and bias control
The first passes used LibriSpeech test-clean (small speaker counts, few sentences). That is fine for wiring scorers and wrappers. It is not enough if someone will argue the benchmark is biased or too narrow.
Final numbers come from VCTK (British English multi-speaker corpus; HF mirror badayvedat/VCTK):
| Design choice | Value | Why it matters |
|---|---|---|
| Speakers | 22 (11 female / 11 male) | Gender balance |
| Accents | 11 (English, Scottish, Irish, Northern Irish, Indian, Welsh, Canadian, American, Australian, South African, New Zealand) | Avoids one-accent conclusions |
| Test sentences per speaker | 12 unique held-out texts | Enough text diversity for WER |
| Manifest rows | 264 | One row per speaker × sentence |
| Synth result rows | 1056 | 4 models × 264 |
| Natural rows | 264 | Same-speaker ceiling for SIM |
| Reference clip | ~10–13 s per speaker | Enough signal for cloning / conditioning |
| Leakage | 0 | Test text never appears in the reference transcript; test utt_ids disjoint from ref |
VCTK evaluation design summary: 22 speakers, 11 accents, 12 unique texts per speaker, 1320 scored rows, zero leakage
Figure: Final VCTK sweep design used for all published numbers.
Speaker-level paired bootstrap
scripts/paired_bootstrap.py resamples the 22 speakers with replacement (B=10000, seed 42). Each speaker contributes its 12-sentence mean, so the observed Δ matches the published utterance-level mean difference on this balanced design.
| Metric | Observed Δ (A8 − XTTS) | 95% CI | Excludes 0? |
|---|---|---|---|
| UTMOS | +0.1434 | [+0.0619, +0.2228] | yes |
| SIM | −0.0116 | [−0.0336, +0.0130] | no |
| WER | −0.0028 | [−0.0072, +0.0016] | no |
Same seed reproduces output/stats_bootstrap.json bit-for-bit.
Harness architecture
manifest_vctk.jsonl
│
▼
run_eval.py
├── wrappers/ pocket_tts, kokoro, audio8, xtts (subprocess)
├── energy gate (reject silent / empty audio)
└── scorers/ speaker_sim (ECAPA), naturalness (UTMOS), wer, timing
│
▼
output/results.csv → aggregated_results.csv
→ natural_results.csv
→ zero_shot_comparison.*
→ stats_bootstrap.json
→ report.md
XTTS lives in venv_xtts and is invoked through xtts_subprocess_entry.py so Coqui’s transformers pin does not break the main env. Audio8 registers ONNX voices against a short reference clip plus transcript under voice_name=speaker_id in voices_vctk. Kokoro and Pocket TTS take documented fixed voices when true cloning is unavailable.
Hardware for the published run: CPU only, 8 cores, ~62 GB RAM, no GPU.
Three bugs
1. Repeated VCTK elicitation text
VCTK reuses elicitation sentences across recordings. The first manifest builder took the first N non-reference utterances without deduplicating by transcript text. Result: only 6 unique texts per speaker out of 12 rows. Distinct audio files, but repeated text pairs. WER looked artificially strong on short repeated lines.
Fix: seen_test_text in scripts/build_vctk_manifest.py, candidate pool 30 → 60, separation on both utt_id and text, rebuild + re-run. Post-fix: 12 unique test texts per speaker, 0/264 leakage rows, all 1056 rows status=OK.
With duplicated elicitation text, WER looked too good (for example Kokoro ~0.0017). After unique texts, Kokoro WER moved to 0.0051 and XTTS to 0.0091. The ranking story stayed directionally similar, but the absolute WER numbers became honest for a multi-sentence claim.
2. Untrained WavLM-large speaker-similarity head
The first published SIM used:
WavLMForXVector.from_pretrained("microsoft/wavlm-large")
emb = model(**inputs).embeddings
microsoft/wavlm-large is an SSL backbone. The trained speaker-verification checkpoints are microsoft/wavlm-base-sv and microsoft/wavlm-base-plus-sv. Loading WavLMForXVector on wavlm-large attaches an untrained x-vector head. Cosine in that space is ~0.98 for almost any two clean speech clips.
That is exactly what the table showed: all four models 0.9856–0.9882; Kokoro af_heart vs 11 male speakers 0.9863. A fixed American female voice cannot be 0.986 similar to eleven male British speakers. Calling this “saturation” was the wrong diagnosis.
What we did: did not re-run TTS. Replaced the scorer with ECAPA-TDNN (speechbrain/spkrec-ecapa-voxceleb), ran a fail-closed sanity gate (same-speaker natural 0.7364 vs cross-speaker 0.0766 vs Kokoro-vs-male 0.0521, n=11 male refs × Kokoro sent1), rescored 1056 synth rows and 264 natural rows, rewrote every table from the new CSVs. Broken column kept at output/results_sim_broken_backup.csv. The full 132-row Kokoro-vs-male mean in the report is 0.0529.
3. Audio8 stuck on the first registered voice (then a registration-only re-synth)
After ECAPA made identity readable, the first Audio8 row sat at 0.1580 — next to Pocket TTS 0.1402, not next to XTTS 0.5569. That was not “weak cloning.” The wrapper was a hard bug:
def clone(..., voice_name="pilot_speaker", ...):
_register_voice(..., voice_name) # cache key was (model_dir, name)
runtime.synthesize(..., voice=voice_name)
run_eval.py never passed a per-speaker name. After the first speaker, _REGISTERED short-circuited and later refs were ignored. The voices dir contained exactly one entry: pilot_speaker. Later “speakers” still looked like p226.
Stuck-run check (one sent1 clip per speaker):
| model | n | matched | mismatch | vs first speaker (p226) | gap |
|---|---|---|---|---|---|
| audio8 (stuck) | 22 | 0.1437 | 0.1439 | 0.4265 | −0.0002 |
| xtts | 22 | 0.4896 | 0.1237 | 0.1771 | +0.3659 |
| pocket_tts | 22 | 0.1321 | 0.1527 | 0.1380 | −0.0207 |
| kokoro | 22 | 0.0289 | 0.0402 | 0.0482 | −0.0114 |
XTTS 0.5569 is the full n=264 mean; 0.4896 is the n=22 sent1 mismatch-check mean.
What we did: fixed the wrapper (voice_name=speaker_id, cache key includes the reference SHA-256, overwrite=True, fresh voices_vctk). Re-synthesized Audio8 only on the existing 264-row VCTK manifest — registration-only, not a new corpus, not new sentences, not a model upgrade. Did not re-run XTTS / Kokoro / Pocket. Backup of the stuck run: output/backup_audio8_stuck/.
After the fix: Audio8 sent1 matched 0.4885 vs mismatch 0.1130 (gap +0.3755, SIM varies). Published Audio8 SIM is 0.5454 (near-tie with XTTS 0.5569; 10/22 vs 12/22). The 0.1580 / matched≈mismatch numbers are this postmortem only.
The earlier LibriSpeech “Audio8 led XTTS by a larger UTMOS margin” comparison is contaminated by the same wrapper defect (one reused voice, not a corpus-sensitivity lesson). Do not cite it as evidence that naturalness rankings shift by corpus.
Repo
What Neo actually did
Neo is a fully autonomous AI engineering agent for fine-tuning, evaluation, classical ML, RAG pipelines, and related shipping work. In VS Code or Cursor it takes a high-level goal, plans, writes code, runs commands, and iterates until the job is done.
For this case study, the human side was essentially the goal and follow-up pressure (“make this defensible for people who know TTS”). Neo:
- Designed a multi-model CPU evaluation harness (
run_eval.py, wrappers, scorers) - Integrated four different model stacks, including an isolated XTTS environment
- Started on LibriSpeech for bring-up, then moved to a VCTK multi-speaker, multi-accent design
- Built manifest generation with gender balance, accent coverage, and leakage checks
- Ran the full 1056-row sweep and wrote comparative reports
- When asked to prove scores were real, re-validated scorers (including direct Whisper checks on WAVs)
- Found the VCTK repeated-sentence bug, fixed the builder, re-ran the sweep, and rewrote the methodology
- After review: replaced the untrained WavLM-large SIM scorer with ECAPA, sanity-gated the metrics, rescored 1056 + 264 rows without re-running TTS, and rewrote README / blog / report / figures from the new CSVs
- After a second review: scored existing WAVs matched vs mismatch and found the Audio8 wrapper never re-registered after
pilot_speaker/p226(stuck SIM 0.1580) - Fixed the wrapper (
voice_name=speaker_id, digest cache,voices_vctk), re-synthesized Audio8 only (registration-only) on the existing VCTK manifest, rescored, and rewrote README / blog / report from SIM 0.5454 - Added a speaker-level paired bootstrap (
scripts/paired_bootstrap.py→output/stats_bootstrap.json) and reordered this write-up so results come before postmortems - Packaged README,
.gitignore, Apache-2.0 license, and pushed github.com/gauravvij/voice-clone-eval
You do not need to re-implement that loop by hand to build on it. Clone the repo and drive the next experiment with Neo.
Replicate or extend this with Neo
1. Clone the project
git clone https://github.com/gauravvij/voice-clone-eval.git
cd voice-clone-eval
Follow README.md for the two venvs (./venv, ./venv_xtts), model downloads, and:
./venv/bin/python run_eval.py \
--manifest data/manifest_vctk.jsonl \
--results output/results.csv
To refresh speaker similarity, bootstrap CIs, and figures only (no TTS):
./venv/bin/python scripts/sanity_gate_scorers.py
./venv/bin/python scripts/rescore_speaker_sim.py
./venv/bin/python scripts/check_audio8_registration.py
./venv/bin/python scripts/zero_shot_comparison.py
./venv/bin/python scripts/paired_bootstrap.py
./venv/bin/python scripts/generate_blog_figures.py
Large artifacts (model weights, raw VCTK audio) are gitignored on purpose. Rebuild or download them locally as documented in the README.
2. Open the folder in VS Code or Cursor with Neo
Install the Neo extension, open this repo, and give Neo a concrete next goal. Examples that map cleanly onto the existing harness:
Add a fifth model
Clone https://github.com/gauravvij/voice-clone-eval and add OpenVoice (or another model I name) as a fifth wrapper. Keep the same scorers and VCTK manifest. Re-run evaluation for the new model only if possible, then update aggregated_results.csv, zero-shot tables if it is a true cloner, and report.md.
GPU path and fair latency
Extend this repo so XTTS and Audio8 can run on CUDA when a GPU is present, keep CPU fallback, and log device + RTF side by side. Do not change the VCTK speaker/sentence design.
Stronger intelligibility scoring
Swap faster-whisper small for whisper medium (or large-v3) on CPU or GPU, re-score existing WAVs in output/audio without regenerating speech if files exist, and compare WER deltas in a short appendix in report.md.
Human spot-check set
Sample 40 clips stratified by model and gender, write a simple HTML listening sheet with hidden model labels, and store the sample list in data/listening_sample.jsonl.
New language or corpus
Replace VCTK with a multilingual corpus I specify. Keep leakage controls and unique-text dedupe. Rebuild manifest, run all models that support the language, and regenerate the report.
CI smoke test
Add a GitHub Actions workflow that installs CPU deps, runs a 2-speaker × 2-sentence smoke manifest, and fails if any row is not status=OK.
3. How to prompt Neo well
- Point at the repo path or clone URL first
- State the decision you care about (latency, WER, true cloning only, etc.)
- Say what must stay fixed (VCTK 22×12, metric definitions, Apache license)
- Ask for regenerated artifacts (
results.csv,report.md) not only code edits
Neo’s job is the full loop: plan, implement, run, notice broken assumptions, fix, and leave evidence in the repo.
Bottom line
See the Audio8 stuck-voice postmortem in output/report.md for the first publish (SIM 0.1580).
- Both intended cloners work. Identity is a near-tie: Audio8 0.5454 vs XTTS 0.5569 (10/22 vs 12/22; natural ceiling 0.6964; SIM CI includes 0). Audio8
sent1matched 0.4885 vs mismatch 0.1130 (gap +0.3755). - Among cloners on this CPU, Audio8 is the practical pick — honest RTF 1.4214 vs XTTS 5.3112, UTMOS 3.3657 (19/22; CI excludes 0).
- For CPU fixed-voice quality and speed, Kokoro wins this sweep (UTMOS 3.6899, RTF 0.1238).
- Do not lead with WER. Natural 0.0355 vs synth 0.0051–0.0091 is Whisper rewarding hyper-articulation. WER CI includes 0.
- Do not cite the old ~0.98 SIM column. It was an untrained head, not a saturated metric.
- The evaluation only became trustworthy after unique-text dedupe on VCTK, a working speaker-verification scorer, a matched-vs-mismatch registration check, a registration-only Audio8 re-synth, and a speaker-level paired bootstrap.
If you want the raw tables, methodology, and code path Neo left behind, start here: https://github.com/gauravvij/voice-clone-eval.
Changelog
- 2026-08-15 — v3. Speaker-level paired bootstrap (B=10000) on Audio8 vs XTTS. UTMOS Δ +0.1434, 95% CI [+0.062, +0.223] excludes 0; SIM Δ −0.0116 CI includes 0. Recommendation reversed: Audio8 is no longer “catalog-like / do not use for cloning.” After the registration-only re-synth it is a working cloner and the CPU pick among cloners (identity near-tie with XTTS; faster; UTMOS lead among cloners). Blog reordered Lead → Results → Recommendations → Methodology → Three bugs → Repo. README leads with the four-model decision table.
- 2026-08-15 — v2. Audio8 wrapper fix + registration-only re-synth (264 clips). Published SIM 0.5454 vs XTTS 0.5569. Stuck-voice 0.1580 is postmortem only. XTTS / Kokoro / Pocket not re-run.
- 2026-08-14 — v1. ECAPA rescore of the original 1056 WAVs. Invalid WavLM-large SIM column retired. Audio8 still stuck on
pilot_speaker/ p226 in that publish.
Try NEO in Your IDE
Install the NEO extension to bring AI-powered development directly into your workflow:
- VS Code: NEO in VS Code
- Cursor: Install NEO for Cursor
