Blog

207 articles

Research, case studies, and engineering deep-dives from the Neo team.

Same Audio, Different Engineering: Kimi K3 vs Opus 5
LLM Evaluation & Benchmarking

Same Audio, Different Engineering: Kimi K3 vs Opus 5

Same Kokoro-82M CPU task via NEO BYOK. Audio tie; speed not comparable. Kimi finished (72); Opus designed the stronger benchmark (63).

July 30, 2026·9 min
Read
Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering

Three AutoML artifacts, all green suites (45/45, 83/83, 37/37). Execution found silent trust failures. Scores: Kimi 68, Fable 61, GLM 54.

July 24, 2026·10 min
Read
Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering

Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.

July 20, 2026·8 min
Read
Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC
Model Optimization & Inference

Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC

Neo profiled NVIDIA Parakeet TDT 0.6B v3 on CPU, ran keep/discard ladders, and froze a static-QDQ production pack that cut primary RTF by ~2.07× on EPYC (~1.42× on Apple Silicon). Runtime-only knobs never cleared 5%.

July 14, 2026·14 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark

Kyutai's Pocket TTS joins the CPU TTS benchmark: 6 configs, 180 timed runs, and 36 WAV samples across RTF, latency, throughput, and UTMOS MOS, plus zero-shot voice cloning from 5 seconds of audio. Built end-to-end with Neo.

July 6, 2026·16 min
Read
Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval
LLM Evaluation & Benchmarking

Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval

Neo evaluated Qwythos-9B at Q4_K_M and Q8_0 on GSM8K, IFEval, and HumanEval from a single prompt. GSM8K hit 84%, IFEval 66%, HumanEval 0% — and Q4 is nearly as good as Q8 for math.

July 3, 2026·10 min
Read
Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.
LLM Evaluation & Benchmarking

Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.

Claude Sonnet vs Ornith:35b on CodeArena — an AI coding benchmark where both passed 7/24 tests. Process scores diverged sharply: self-finalization, tool mix, and $0 local cost vs $2.73 API fees.

July 2, 2026·14 min
Read
NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously
LLM Evaluation & Benchmarking

NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously

NEO built and ran the Ornith Evaluation Framework autonomously on Ornith-1.0-35B: 100/100 terminal safety, Level 6/15 skill ceiling. What the model scored and how the harness was verified.

June 27, 2026·9 min
Read
GLM 5.2 Built TrackLab: Browser Computer Vision with NEO BYOK
LLM Evaluation & Benchmarking

GLM 5.2 Built TrackLab: Browser Computer Vision with NEO BYOK

GLM 5.2 built TrackLab end to end — a browser CV studio with detection, tracking, and line counting — entirely through NEO BYOK. Same agent workflow, different model.

June 26, 2026·10 min
Read
GLM 5.2 vs Kimi K2.6: Same Agent Workflow, Same Citation Bug, Opposite Failures
LLM Evaluation & Benchmarking

GLM 5.2 vs Kimi K2.6: Same Agent Workflow, Same Citation Bug, Opposite Failures

What NEO's BYOK found comparing GLM 5.2 vs Kimi K2.6 in VS Code: Kimi won capability, GLM craftsmanship. Both failed Article 75 differently. Specs and setup guide.

June 23, 2026·8 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1: A Real CPU TTS Benchmark

A CPU-only TTS benchmark of Kokoro 82M, Supertonic 3, and Inflect-Nano-v1 across RTF, latency, throughput, and UTMOS MOS: 5 configs, 150 timed runs, and 30 WAV samples. Built end-to-end with Neo.

June 22, 2026·14 min
Read
From Synthetic Data Generation to Dataset Engineering: What Changed When We Added Neo MCP
LLM Evaluation & Benchmarking

From Synthetic Data Generation to Dataset Engineering: What Changed When We Added Neo MCP

874 real-world agent failure records from a 7-phase governed pipeline: one prompt, HuggingFace ingestion, adversarial verification, and full provenance without synthetic padding.

June 17, 2026·14 min
Read