Open to Incident Commander & SRE roles

Calm in the chaos.
Automation that kills the toil.

I'm a Senior SRE and Incident Commander with 15+ years on the front line. EBay-scale e-commerce, high-growth fintech, multi-cloud platforms. Today I build the AI-native version of that job: a security operations centre that triages itself, a NOC that heals itself, and incident response measured in minutes instead of meetings.

Las Vegas, NV · remote
Portrait of Jeremy Martinez, Senior Site Reliability Engineer and Incident Commander
15+
years in production operations
99.997%
uptime SLA held at eBay
90%
of a team's manual toil automated away
0
postmortem findings left as prose

The short version

I spent fifteen years doing operations by hand. Now I build the thing that does it.

Most operations teams are not short on tooling. They are short on time, because the same fifty alerts get triaged by hand every day, the same runbook gets re-read at 3am, and the same postmortem action item quietly stops being true six weeks later.

So I built a fleet where the machine does that work and the human keeps the judgment calls, across the whole arc of an incident. Detection stops being a firehose and becomes a short list with the evidence already attached. Diagnosis starts at understanding instead of starting at collection, which is where the clock actually goes. Response happens without waiting for somebody to be paged, wake up, and find their laptop, and that wait is the single largest number in mean time to recovery that no amount of hiring removes. Recovery gets drilled instead of assumed. And every lesson from every postmortem becomes an assertion that re-proves itself on a schedule, because an action item nobody checks again is just a lie with a due date.

None of that is a demo. It runs continuously, on hosts carrying tax records and health data, because testing against the strictest rules you can find is the honest way to learn whether something holds. It is not where the work stops applying. An outage does not care what industry the data belongs to, and this is the same discipline I brought to incident command at eBay and Upstart, with the toil removed.

Read the deep dives

Incident command

IC for global outages at eBay, Upstart and Dynascale. Severity triage, escalation, exec comms under pressure, and blameless postmortems that produce corrective actions instead of theater.

AI-native operations

An AI SOC that triages every alert, an autonomous response engine with a constrained action space, and a headless responder that resolves incidents or escalates with evidence.

Detection & response

SIEM, host IDS, kernel audit, eBPF runtime and Suricata NDR. Deployed in stages, each one verified with a test designed so that it can fail.

Toil elimination

Took 90% of the manual operational work off a 12-person team at eBay, so they could spend the week on engineering instead. The same instinct now runs a multi-tenant fleet on continuous invariant assertion instead of a checklist.

Platform automation

Terraform, Ansible, Kubernetes, Argo CD, Jenkins, plus self-healing remediation, automated patching with canaries, and restore drills that export metrics rather than reassurance.

Regulated environments

Off-host key custody, per-record authorization, audit logging, and change control an auditor can follow. Built against GLBA Safeguards, IRS Pub 4557, ESIGN/UETA and PHI-adjacent care, because the strictest bar is the honest one to test against.

Selected systems

Things I've built.

All thirteen deep dives →

Looking for someone to own incidents. Or to modernize how you handle them.

Incident Commander, Senior/Staff SRE, or the person who builds the AI operations layer your team keeps talking about. I'm also available for keynotes, workshops, and internal training.