Blog

210 articles

Research, case studies, and engineering deep-dives from the Neo team.

Grok 4.6 vs Kimi K3 vs Opus 5 on Kokoro-82M TTS: Strong Research, No Production-Ready Speedup
LLM Evaluation & Benchmarking

Grok 4.6 vs Kimi K3 vs Opus 5 on Kokoro-82M TTS: Strong Research, No Production-Ready Speedup

Grok 4.6, Kimi K3, and Opus 5 each tried to speed up Kokoro-82M TTS on CPU. Study scores: Opus 83, Kimi 60, Grok 51. Grok's best verified gain was about 6–8% in one setup, but audio checks failed, so nothing shipped.

August 13, 2026·8 min
Read
Making OmniVoice 6.2× Faster Than Its 32-Step Default Without Retraining
Model Optimization & Inference

Making OmniVoice 6.2× Faster Than Its 32-Step Default Without Retraining

OmniVoice sounded good but felt slow. Without training a new model, we made it about 6× faster than the default path, and still about 3× faster than the official fast setting, then kept the quality checks honest.

August 12, 2026·15 min
Read
Three ASR models tied on the benchmark. One was almost three times worse on real speech.
LLM Evaluation & Benchmarking

Three ASR models tied on the benchmark. One was almost three times worse on real speech.

Five ASR models on 296 real clips across 19 conditions. Clean-speech WER ties at 4.4%, then accent and noise separate the field, with committed transcripts, bootstrap CIs, and a CI regression gate.

August 11, 2026·16 min
Read
Evaluating Voice Cloning Models on CPU: A Practical Benchmark of Pocket TTS, Kokoro, Audio8, and XTTS-v2
LLM Evaluation & Benchmarking

Evaluating Voice Cloning Models on CPU: A Practical Benchmark of Pocket TTS, Kokoro, Audio8, and XTTS-v2

A CPU-only VCTK sweep of Pocket TTS, Kokoro, Audio8, and XTTS-v2 across speaker similarity, UTMOS, Whisper WER, and RTF — with a true zero-shot Audio8 vs XTTS head-to-head. Built end-to-end with Neo.

August 7, 2026·12 min
Read
Kimi K3 vs Opus 5: Which Model Built the Better Kokoro CPU Optimizer?
LLM Evaluation & Benchmarking

Kimi K3 vs Opus 5: Which Model Built the Better Kokoro CPU Optimizer?

Using Neo BYOK, both models optimized Kokoro-82M on CPU. Opus built the stronger study, but its claimed 18.66% speedup became 9.81% slower on replay.

August 1, 2026·8 min
Read
Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering

Three AutoML artifacts, all green suites (45/45, 83/83, 37/37). Execution found silent trust failures. Scores: Kimi 68, Fable 61, GLM 54.

July 24, 2026·10 min
Read
Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering

Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.

July 20, 2026·8 min
Read
Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC
Model Optimization & Inference

Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC

Neo profiled NVIDIA Parakeet TDT 0.6B v3 on CPU, ran keep/discard ladders, and froze a static-QDQ production pack that cut primary RTF by ~2.07× on EPYC (~1.42× on Apple Silicon). Runtime-only knobs never cleared 5%.

July 14, 2026·14 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark

Kyutai's Pocket TTS joins the CPU TTS benchmark: 6 configs, 180 timed runs, and 36 WAV samples across RTF, latency, throughput, and UTMOS MOS, plus zero-shot voice cloning from 5 seconds of audio. Built end-to-end with Neo.

July 6, 2026·16 min
Read
Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval
LLM Evaluation & Benchmarking

Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval

Neo evaluated Qwythos-9B at Q4_K_M and Q8_0 on GSM8K, IFEval, and HumanEval from a single prompt. GSM8K hit 84%, IFEval 66%, HumanEval 0% — and Q4 is nearly as good as Q8 for math.

July 3, 2026·10 min
Read
Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.
LLM Evaluation & Benchmarking

Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.

Claude Sonnet vs Ornith:35b on CodeArena — an AI coding benchmark where both passed 7/24 tests. Process scores diverged sharply: self-finalization, tool mix, and $0 local cost vs $2.73 API fees.

July 2, 2026·14 min
Read
NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously
LLM Evaluation & Benchmarking

NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously

NEO built and ran the Ornith Evaluation Framework autonomously on Ornith-1.0-35B: 100/100 terminal safety, Level 6/15 skill ceiling. What the model scored and how the harness was verified.

June 27, 2026·9 min
Read