Linux
Infrastructure
The foundation everything else stands on: production Linux since 2012 — from 15,000-instance fleets to hardened modern infrastructure, and from firefighting to automation-first operations.
What I mean by it
I started here. Every later layer — cloud, platform, agents — is easier to reason about when you've personally watched a misconfigured sysctl take down a revenue system at scale.
The documented fleet record
- ▸ Yesmail (2012–2014). Stabilized a troubled environment and reduced SEV 1 alerts to near-zero; automated deployments with Shell and Chef.
- ▸ Intuit (2014–2015). Decommissioned bare-metal and virtual resources inside a 5,000+ instance production network.
- ▸ Epsilon (2015–2016). Automated on-demand environment provisioning with Ansible and Puppet across 15,000+ AWS and VMware instances.
- ▸ OCP (2020–2021). Bash/Ansible automation of Linux environments; AWS provisioning via Terraform; MySQL/MongoDB and VMware vSphere.
- ▸ ConstructConnect (2021–2025). Automated Linux patching with Ansible/GitLab at fleet scale before moving up to the platform layer.
Employer-stated role outcomes, Sep 2012–2025 · RHEL, CentOS, Ubuntu, Bash, Ansible, Chef, Puppet
Authority, checked by others
Since 2020, Packt Publishing has paid me to technically review Linux and infrastructure manuscripts before publication — most notably both editions of "Mastering Linux Administration" (2021, 2024). Reviewing a 500-page administration book means every command, flag and architectural claim in it has to survive someone who runs these systems for a living.
On the tooling side I write Linux software, not just operate it: NetRain (real-time Rust packet monitor with DDoS/port-scan detection) and k8s-netinspect (Kubernetes network inspection).
Common questions
Why does Linux depth matter in the AI era?
Because AI systems run on Linux, and their failure modes surface as Linux problems first — memory exhaustion, socket leaks, DNS, permissions. Agents that deploy software need an operator who can read what the OS is actually doing. My agentic harness runs on exactly this foundation.
What's your automation approach for an existing fleet?
Measure first, automate the painful repeatable thing, keep every change in version control, and make rollback boring. Ansible for declared state, GitLab pipelines for sequencing, and never a manual step that outlives the incident that justified it.