Fixing logrotate's Glob Confusion

A follow-up to the wildcard bug post from last year. Picking the Thread Back Up The rename workaround, *.example.com to _.example.com, fully resolved the operational problem. Retention became predictable again, and there was no pressing reason to go further. But a workaround isn’t a fix, and the actual bug was still sitting in logrotate’s own source, waiting to bite the next person who names a file after a wildcard domain. This week I finally sat down and found it properly, instead of just renaming files around it. ...

August 20, 2026 · Alexander Bakin

Building a Production-Grade Personal Blog with AWS, Terraform, and Hugo

Motivation As a DevOps engineer, your personal website is your portfolio. It should demonstrate not just what you know but what you build. This blog is itself an example of a production-grade cloud architecture - every decision documented, every tradeoff explained. Architecture Overview The stack: Layer Technology Content Markdown → Hugo static site CI/CD (Infra) GitHub Actions - PR plan, merge apply CI/CD (Content) GitHub Actions - build, sync, invalidate State Mgmt S3 backend, concurrency-gated at pipeline level Origin S3 (private, versioned, encrypted, CloudFront OAC only) CDN CloudFront with HTTPS, Brotli/Gzip, security headers DNS Route 53 alias records (A / AAAA, root + www) TLS ACM certificate (auto-renewal, TLSv1.2_2021) Auth OIDC - no AWS access keys stored anywhere Monitoring CloudWatch dashboard + error rate alarm Cost Control AWS Budgets alert (direct email) Key Design Decisions 1. Two Separate Pipelines Infrastructure changes and content changes have different risk profiles and review requirements. They’re handled by separate workflows: ...

August 14, 2026 · Alexander Bakin

Getting to 99.9%: Kubernetes and Real Observability for 20+ Microservices

99.9% availability sounds like a slogan until you do the arithmetic: it’s a budget of roughly eight and a half hours of downtime for an entire year, across every service on the platform. Once you frame it that way, it stops being an aspiration and becomes an engineering constraint you either design for or blow through by month three. Getting there for a platform of 20+ microservices took deliberate work in two places that don’t get enough credit next to “just add more replicas”: the rollout path, and observability that’s actually fast enough to matter. ...

September 15, 2025 · Alexander Bakin

The Wildcard That Broke Nginx Log Rotation

One of those infrastructure problems that looks simple, until you realize the filename itself is part of the problem. The Setup We had Nginx configured to handle wildcard subdomains: server_name *.example.com;, with access logs named after the server block. That’s a convenient pattern for multi-tenant setups where subdomains get created dynamically and you don’t want to hand-maintain a vhost file per tenant. Nginx doesn’t care that the resulting log filename contains a literal * - it just opens the file and writes to it. ...

July 5, 2025 · Alexander Bakin

From 4 Hours to 20 Minutes: Replacing a Manual Deployment with GitLab CI

Deploying used to mean an engineer, a runbook document, and four hours of undivided attention. Not four hours of a pipeline running in the background, four hours of a person actively driving it: SSH into the target host, pull the latest code, run the build script by hand, stop the old process, start the new one, watch the logs to see if it came up clean, and post a status update. It worked, in the sense that it got code to production. It didn’t scale, in the sense that “get code to production” now depended entirely on one specific person having a free afternoon. ...

April 9, 2024 · Alexander Bakin

Writing a Prometheus Exporter in Go for Metrics Nobody Was Collecting

The platform I operate is a fleet of NGINX instances running a Web Application Firewall module in front of production traffic. The stock NGINX exporter gave us request counts, status codes, and upstream latency, all useful, and all completely blind to the one question that actually mattered for this platform: is the WAF itself doing its job, or is it about to fall over. The Gap The WAF module exposes its own internal counters through a local status endpoint: rule hits, rule blocks broken down by rule ID, requests evaluated per second, memory used by the rule engine. None of it reaches Prometheus by default, because it’s not something the generic NGINX exporter knows exists. Our dashboards could tell us “NGINX is up.” They couldn’t tell us “rule 942100 just started blocking forty times its normal traffic,” which is exactly the kind of thing that means either an attack is underway or a rule update just went sideways and started blocking legitimate users. Both are pages. Neither showed up anywhere. ...

August 21, 2023 · Alexander Bakin

Exposing a Homelab to the Internet Without Exposing Myself

Once the homelab had more than one service worth using away from home (Nextcloud for files, Jellyfin for media), “just port forward it” stopped being an option I was willing to consider. I spend enough of my working hours dealing with what happens to services that sit on the open internet unpatched or misconfigured for a week. I wasn’t going to do that to my own network. Layer One: A Private Network That Doesn’t Need Open Ports Tailscale solved the “I want access from anywhere” problem without opening a single inbound port on my router. It’s WireGuard under the hood: every device gets a key, joins the same private tailnet, and talks to every other device directly (or through a relay when direct connection isn’t possible) over an encrypted tunnel. My phone, laptop, and homelab nodes all land on the same private address space no matter which network they’re actually sitting on. For anything only I need to reach, that’s the entire solution: no public DNS record, no exposed port, nothing for an internet-wide scanner to ever find. ...

June 12, 2022 · Alexander Bakin

Deploying Vaultwarden with Ansible: My Homelab's First Real Service

I started self-hosting for a boring reason: I didn’t like that a single cloud account breach could hand someone every password I own. A self-hosted password manager has its own risks, but at least the blast radius is mine to control. Vaultwarden (a lightweight Bitwarden-compatible server) was the obvious first service, and I used it as an excuse to do my homelab properly instead of SSH-ing in and running commands by hand. ...

July 18, 2021 · Alexander Bakin