Our point of view on shipping AI — page 4 of 7.
- Dev Tools6 min
Your Refresh Token Is Now a Lock, Not a Secret
By Petru Popa ·OAuth apps on GitHub can now register up to ten redirect URIs and refresh short-lived tokens. The authorization docs add the part the changelog leaves out: refreshing invalidates the refresh token and the old access token together, which turns a copyable secret into something exactly one process may hold.
Read - AI Engineering6 min
LLM Evaluation Is Three Different Jobs
By Petru Popa ·Picking a model, catching a regression, and watching production are three different measurement problems that share one name. Teams get into trouble when the artifact from one gets used to answer another — and the LLM-as-judge shortcut has failure modes with numbers attached.
Read - AI Engineering6 min
LLM Observability Is Mostly a Schema Problem
By Petru Popa ·Buying a dashboard is the easy half. The hard half is whether your traces carry the fields you will need at two in the morning — and the standard that defines those fields is still marked as in development.
Read - AI Engineering6 min
Pick LLM Evaluation Metrics by What You Can Supply
By Petru Popa ·The usual taxonomy sorts metrics by how they work — n-gram, embedding, model judge. That axis is useless for choosing. Sort them instead by what reference data each one demands from you, because that is the constraint you actually have.
Read - AI Engineering6 min
LLM Testing Already Had a Playbook in 2020
By Petru Popa ·Teams treat testing a language model as a new discipline with no precedent. The methodology that fits it best was published in 2020, and the test type it names as most valuable is the one almost nobody writes.
Read - AI Engineering6 min
Fine Tuning vs RAG Is the Wrong Question
By Petru Popa ·Two papers settle the versus framing between them. One finds retrieval beats unsupervised fine-tuning at teaching a model facts. The other finds the gains from both are cumulative — and things that add are not substitutes.
Read - AI Strategy6 min
The Watermark Is Sparsest Where You Would Audit
By Petru Popa ·Anthropic will watermark future Claude models using a version of SynthID-Text. The announcement and the Nature paper behind it agree that the mark weakens on code, short answers, and factual passages, which is where most enterprise detection requirements are aimed.
Read - AI Engineering6 min
Verified Is Not Reproduced
By Petru Popa ·A community effort reproduced claims from 2,226 ICML 2026 papers and reported 51% with at least one verified claim. Opening the frozen verdicts dataset shows what that percentage covers, what a $20 compute cap excludes, and why there is no way to look your paper up.
Read - Dev Tools6 min
Conformance Lets a Client Ignore Half Your Plugin
By Petru Popa ·GitHub shipped Agent Plugins 1.0 across VS Code, Copilot CLI, the Copilot SDK and the Copilot app. The spec repository behind it standardizes the package and explicitly defers trust, permissions, sandboxing and provenance to the client — and lets a conforming client implement skills without MCP servers.
Read - Dev Tools6 min
The Audience Is the One Claim Your Attacker Picks
By Petru Popa ·A proposal to let GitHub Actions workflows pre-declare OIDC audiences is the wrong fix to wait for. Reading the permission schema and AWS's condition-key reference shows why the control that holds today lives on the relying party.
Read - AI News6 min
The 24 GB Card Fits the Weights, Not the Features
By Petru Popa ·Meta's Muse Glimmer targets consumer GPUs with 24 to 32 GB. Reading config.json and the GGUF file listing, the binding constraint on a 24 GB card is not the 131,072-token context — it is the vision projector and the speculative-decoding drafter, the two files that are the reason to pick this model.
Read - AI Engineering5 min
The 3.3x Belongs to the Model, Not the Flag
By Petru Popa ·vLLM published a 3.3x throughput gain for decode context parallelism on Kimi K2.6. The config source shows the ceiling for grouped-query models is set by KV head count — and for Qwen3-235B-A22B on one node, that ceiling is two.
Read