Open to Incident Commander & SRE roles
Calm in the chaos.
Automation that kills the toil.
I'm a Senior SRE and Incident Commander with 15+ years on the front line. EBay-scale e-commerce, high-growth fintech, multi-cloud platforms. Today I build the AI-native version of that job: a security operations centre that triages itself, a NOC that heals itself, and incident response measured in minutes instead of meetings.

- 15+
- years in production operations
- 99.997%
- uptime SLA held at eBay
- 90%
- of a team's manual toil automated away
- 0
- postmortem findings left as prose
The short version
I spent fifteen years doing operations by hand. Now I build the thing that does it.
Most operations teams are not short on tooling. They are short on time, because the same fifty alerts get triaged by hand every day, the same runbook gets re-read at 3am, and the same postmortem action item quietly stops being true six weeks later.
So I built a fleet where the machine does that work and the human keeps the judgment calls, across the whole arc of an incident. Detection stops being a firehose and becomes a short list with the evidence already attached. Diagnosis starts at understanding instead of starting at collection, which is where the clock actually goes. Response happens without waiting for somebody to be paged, wake up, and find their laptop, and that wait is the single largest number in mean time to recovery that no amount of hiring removes. Recovery gets drilled instead of assumed. And every lesson from every postmortem becomes an assertion that re-proves itself on a schedule, because an action item nobody checks again is just a lie with a due date.
None of that is a demo. It runs continuously, on hosts carrying tax records and health data, because testing against the strictest rules you can find is the honest way to learn whether something holds. It is not where the work stops applying. An outage does not care what industry the data belongs to, and this is the same discipline I brought to incident command at eBay and Upstart, with the toil removed.
Read the deep divesIncident command
IC for global outages at eBay, Upstart and Dynascale. Severity triage, escalation, exec comms under pressure, and blameless postmortems that produce corrective actions instead of theater.
AI-native operations
An AI SOC that triages every alert, an autonomous response engine with a constrained action space, and a headless responder that resolves incidents or escalates with evidence.
Detection & response
SIEM, host IDS, kernel audit, eBPF runtime and Suricata NDR. Deployed in stages, each one verified with a test designed so that it can fail.
Toil elimination
Took 90% of the manual operational work off a 12-person team at eBay, so they could spend the week on engineering instead. The same instinct now runs a multi-tenant fleet on continuous invariant assertion instead of a checklist.
Platform automation
Terraform, Ansible, Kubernetes, Argo CD, Jenkins, plus self-healing remediation, automated patching with canaries, and restore drills that export metrics rather than reassurance.
Regulated environments
Off-host key custody, per-record authorization, audit logging, and change control an auditor can follow. Built against GLBA Safeguards, IRS Pub 4557, ESIGN/UETA and PHI-adjacent care, because the strictest bar is the honest one to test against.
Selected systems
Things I've built.
Security & Response
The AI SOC
Everybody has detection. Nobody has anybody to read it. That is not a problem you hire your way out of, and it is not a problem you buy another product to fix.
Security & Response
Autonomous Response
The AI never writes a command. It picks a number off a list, so the worst a hostile input can do is make it pick a different approved action.
Reliability & Automation
The Self-Healing NOC & War Room
A 200 means something is listening. It does not mean the site is up. That gap is where an outage sits quietly while every dashboard stays green.
Reliability & Automation
The Terminal Brain
A committee of models reads over your shoulder in a live production shell. The watching tier cannot type. Not switched off, not configured down. There is no code path that types.
Looking for someone to own incidents. Or to modernize how you handle them.
Incident Commander, Senior/Staff SRE, or the person who builds the AI operations layer your team keeps talking about. I'm also available for keynotes, workshops, and internal training.