AI Engineering — page 1 of 3
- AI Engineering6 min
To Run an LLM Locally in Transformers, the Metal Kernel Has to Be There
By Petru Popa ·Hugging Face shipped a path that runs llama.cpp quants inside transformers without dequantizing them. The packed path needs a Hub kernel whose only published driver family is metal, it forces float32, it covers Qwen3.5 and nothing else, and the fallback when the kernel is absent costs you the whole quantization.
Read - AI Engineering6 min
DSpark Speculative Decoding: The 4x Was Measured on Math, and Tool Calls Accept About Half as Many Draft Tokens
By Petru Popa ·The vLLM team's DSpark drafter for Kimi K3 quadruples single-stream decode speed on math reasoning. The Hugging Face cards for the four published DSpark drafters show tool-calling prompts accept 43 to 59 percent as many extra tokens per step as math does, every published speed chart is a math chart, and vLLM's confidence-scheduled verification, the half of DSpark built for high concurrency, defaults to off.
Read - AI Engineering6 min
The Fastest Local LLM on a Mac Ships From a Repository Two Days Old
By Petru Popa ·Splash is an Apache-2.0 inference engine for Apple silicon that beats every engine Inco AI measured against it, on one machine, on two models. The repository behind it was created on 18 September and has twelve commits. Here is what the issue tracker and the model packages say about adopting it.
Read - AI Engineering5 min
Chord's INT4 MoE Kernel: Up to 2.15x per Layer, 4 to 8 Percent for a Self-Hosted LLM
By Petru Popa ·vLLM and Novita AI announced Chord, an INT4 mixture-of-experts kernel for Kimi K2.x, with per-layer speedups up to 2.15x. The end-to-end run in the same post shows 4.1 to 8.0 percent more decode throughput on eight H200s, the baseline is a Humming release from July, and the install gives vLLM's humming dependency a second owner.
Read - AI Engineering6 min
WeKnora Is Enterprise RAG. Seven of Its Ten Security Advisories Hit the Agent.
By Petru Popa ·Tencent's WeKnora is an open-source enterprise RAG platform trending this week. Seven of its ten security advisories sit in agent tools and none in retrieval, while the quick-start .env file still ships fixed encryption keys with registration open.
Read - AI Engineering6 min
Self-Hosted AI Now Keeps Its KV Cache on Disk, and vLLM Never Expires It
By Petru Popa ·vLLM's tiered KV cache offloading writes prompt-derived blocks to a filesystem or an S3 bucket. The v0.29.0 tier code has no age or capacity expiry, creates block files mode 0644, and names them with a hash any instance can recompute.
Read - AI Engineering6 min
PatchTST-FM-r2 Beats r1 on Average and Loses on 24 of 97 GIFT-Eval Configurations
By Petru Popa ·IBM's Granite PatchTST-FM-r2 ranks second among zero-shot, replicable models on GIFT-Eval, and a recompute from the benchmark's own result files confirms it. Those files also show r2 scoring worse than r1 on a quarter of the benchmark, and the Hub now tags the model OpenMDW rather than Apache.
Read - AI Engineering6 min
AI Agent Governance Does Not Survive The Approval Prompt
By Petru Popa ·GitHub made enterprise managed permissions for Copilot agent operations generally available on 9 September, and the reference documentation says more than the announcement did: the permission keys do not reach the cloud agent at all, and one allow entry turns every unlisted operation into a prompt.
Read - AI Engineering6 min
GHES 3.22 Points Copilot CLI At A Private LLM, One Per Instance
By Petru Popa ·GitHub Enterprise Server 3.22 lets an administrator configure a model provider once so Copilot CLI works without GitHub Cloud. The admin docs show it is a proxy with a single model, and that the isolation comes from the endpoint you pick, not from the feature.
Read - AI Engineering6 min
Ai2 Audited 16 LLM Benchmarks. Nearly Half the Safety Questions Score Reasoning.
By Petru Popa ·Ai2 published a per-question fit of 16 LLM benchmarks and shipped the item table as a dataset. Grouping its 29,574 rows by benchmark shows that BBQ and WMDP, both filed under safety, discriminate on the reasoning dimension instead.
Read - AI Engineering6 min
A Self-Hosted LLM Has One Context Pool and Four Users By Default
By Petru Popa ·llama.cpp v0.4.0 adds a per-slot context limit to its server. Reading the server sources shows the cap is a minimum over three numbers, and that omitting -c and writing -c 0 no longer mean the same thing.
Read - AI Engineering6 min
Hugging Face Shipped 207 WebGPU Kernels. Your Lockfile Pins None of Them.
By Petru Popa ·The npm package is one preview version old. The 207 kernel repositories it loads at runtime carry no git tags, and 63 of them moved on day one. The speed number is the least interesting thing here.
Read