System One Models: Jev, the Open-Source Alternatives, and Where They Fit

Most LLM calls in a backend are structured decisions dressed up as chat. TypeSafe's Jev answers only that kind of question: typed answers with probabilities, in one pass, with no text generation. What it is, how its confidence numbers work, what 'can't hallucinate' does and does not mean, the open-source alternatives that appeared within weeks, and where an SRE would put one.

October 2026

The Layer Under Kubernetes: What a K8s-First SRE Had to Relearn About VMs and Clouds

Which of my Kubernetes-shaped SRE instincts are about the abstraction and which are about the machine underneath, worked out against a hosting company that runs about 12,000 single-tenant VMs across ten cloud providers with no Kubernetes at all.

September 2026

VictoriaMetrics, VictoriaLogs and VictoriaTraces: What Prometheus Left Out, and the Stack Built Around It

VictoriaMetrics started as a remote storage backend for Prometheus and grew into a full monitoring stack. What Prometheus deliberately does not do, why two HA Prometheus replicas are not a cluster, how VictoriaMetrics splits Prometheus's jobs into components and keeps long-term data without object storage, how VictoriaLogs compares with Loki and VictoriaTraces with Tempo and Jaeger, what else the project ships, and when it is the right choice.

October 2026

What I Learned Reading All 71 Articles of the Cyber Resilience Act

I read the whole of Regulation (EU) 2024/2847, every article and every annex, and mapped KubeAid and LinuxAid against it for my employer. These are the things that surprised me, the mental model I ended up with, and what I would tell the next engineer who has to do the same reading.

September 2026

Where Linux Meets Kubernetes: Six Mental Models for When top Isn't Enough

Processes and signals, namespaces and cgroups, memory, storage, networking, and identity: the mechanism behind each, the two or three ways it breaks under a pod, and the one check that confirms it. Written from on-call across 100+ production Kubernetes clusters.

September 2026

CiliumNetworkPolicy Anti-Patterns: 8 Ways Correct-Looking YAML Silently Drops Traffic

Lessons from writing CiliumNetworkPolicies for the KubeAid-addons chart - toServices port traps, labels Cilium ignores, toFQDNs without a DNS rule, default-deny surprises, and why every namespace needs exactly one default-deny.

August 2026

Ten Things to Get Right Before You Run Loki in Production

Deployment mode, object storage, label cardinality, multi-tenancy and auth, cross-cluster shipping, and the failures that are silent rather than loud. Written after moving a Loki install twice and debugging an object store that corrupted everything it was given.

August 2026

Why Your DaemonSet Rolling Update Gets Stuck on hostPort Pods

A breakdown of why DaemonSets using hostPort can get stuck in Pending during rolling updates, and why the default maxSurge/maxUnavailable settings aren't always right for them.

July 2026

MCP Explained: What It Is and Why It Exists

A breakdown of the Model Context Protocol - the problem it solves, its architecture, and how a tool call actually works under the hood.

July 2026

Pod Disruption Budgets Explained: How Kubernetes Protects You From Voluntary Disruptions

What a PodDisruptionBudget actually does, the difference between voluntary and involuntary disruptions, how the eviction API respects it, and the mistakes that quietly break node drains and cluster upgrades.

July 2026