<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Prometheus on Alexander's Blog</title><link>https://alexanderbakin.com/tags/prometheus/</link><description>Recent content in Prometheus on Alexander's Blog</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 20 Aug 2026 18:13:44 +0400</lastBuildDate><atom:link href="https://alexanderbakin.com/tags/prometheus/index.xml" rel="self" type="application/rss+xml"/><item><title>Getting to 99.9%: Kubernetes and Real Observability for 20+ Microservices</title><link>https://alexanderbakin.com/k8s-observability/</link><pubDate>Mon, 15 Sep 2025 00:00:00 +0000</pubDate><guid>https://alexanderbakin.com/k8s-observability/</guid><description>&lt;p&gt;99.9% availability sounds like a slogan until you do the arithmetic: it&amp;rsquo;s a budget of roughly eight and a half hours of downtime for an entire year, across every service on the platform. Once you frame it that way, it stops being an aspiration and becomes an engineering constraint you either design for or blow through by month three. Getting there for a platform of 20+ microservices took deliberate work in two places that don&amp;rsquo;t get enough credit next to &amp;ldquo;just add more replicas&amp;rdquo;: the rollout path, and observability that&amp;rsquo;s actually fast enough to matter.&lt;/p&gt;</description></item><item><title>Writing a Prometheus Exporter in Go for Metrics Nobody Was Collecting</title><link>https://alexanderbakin.com/prometheus-exporter/</link><pubDate>Mon, 21 Aug 2023 00:00:00 +0000</pubDate><guid>https://alexanderbakin.com/prometheus-exporter/</guid><description>&lt;p&gt;The platform I operate is a fleet of NGINX instances running a Web Application Firewall module in front of production traffic. The stock NGINX exporter gave us request counts, status codes, and upstream latency, all useful, and all completely blind to the one question that actually mattered for this platform: is the WAF itself doing its job, or is it about to fall over.&lt;/p&gt;
&lt;h2 id="the-gap"&gt;The Gap&lt;/h2&gt;
&lt;p&gt;The WAF module exposes its own internal counters through a local status endpoint: rule hits, rule blocks broken down by rule ID, requests evaluated per second, memory used by the rule engine. None of it reaches Prometheus by default, because it&amp;rsquo;s not something the generic NGINX exporter knows exists. Our dashboards could tell us &amp;ldquo;NGINX is up.&amp;rdquo; They couldn&amp;rsquo;t tell us &amp;ldquo;rule 942100 just started blocking forty times its normal traffic,&amp;rdquo; which is exactly the kind of thing that means either an attack is underway or a rule update just went sideways and started blocking legitimate users. Both are pages. Neither showed up anywhere.&lt;/p&gt;</description></item></channel></rss>