Why SRE?
To me, a reliable system is a product that keeps delivering value to the people who use it. Availability and latency matter to the extent that they show this. What draws me to SRE is the rationality it brings to engineering culture and to product decisions, with agreed criteria in place of opinions and urgency.
Reliability from the user’s point of view
An SLO is only useful when it is tied to the reality of the business. A target chosen because the number looks good, with no relation to what people expect from the service, becomes a number to hit that says nothing about the experience of the people using it. I prefer to start by asking what the people who use a service need from it, and what it costs to fall short.
SRE’s approach to risk starts from a simple idea: no service needs to be 100% reliable, and every extra point of reliability costs time and money and slows down change. With an SLO agreed with product, “the system feels unstable” becomes a question you can answer: how much of the error budget have we used? If there is room left, the team can speed up releases and experiments. If the room is gone, stability becomes the priority, and that was agreed before the incident.
When we decide to accept a risk, I make the uncertainties clear and agree on who is accountable for the decision and when we will revisit it, with an analysis sized to the problem. My years practicing law got me used to this kind of decision, where someone chooses knowing what they are accepting.
Security joins the conversation from the design stage, because security and reliability have a lot in common. Alongside availability and latency, I look at data protection, access control, and potential abuse, and I look for criteria that reflect the consequences for the business and for people. I wrote more about this in Embracing Risk.
Failure as learning
Failure is part of operating software, and I see each failure as a step in a continuous engineering learning process. When a mistake gets through an engineering process, the useful question is what in the process let it get that far. Blaming people for a process error leaves the process as it was.
That is why I see the postmortem as a necessity. It exists to reduce failures and risks and to increase the team’s operational capacity, as long as the actions that come out of it turn into engineering work: automation, tests, better alerts, or design changes. That way, recurring manual intervention becomes code that can be tested, reviewed, and maintained.
Platforms built with teams
I don’t believe in platforms imposed unilaterally. I prefer treating a platform as a product: understanding who will use it, what problem it solves, and whether it is worth the maintenance effort. Self-service and reusable standards only pay off when people actually adopt them.
The goal is for the people who build software to ship it and fix problems without waiting in a request queue, with context and observability to understand the system and safe ways to change it. Standards and automation are worth it when they reduce cognitive load without hiding the system.
The most sophisticated solution is not always the right one. Every extra component is something someone has to operate and maintain, and in many contexts a platform with less complexity serves teams better. In platform decisions, the technical outcome counts, and so do the conditions of the people who will work with it every day.
AI, a new challenge for SRE
AI-assisted development is a new challenge for anyone responsible for reliability. Teams can produce more code and try more ideas, which increases the volume of changes reaching production. To keep up with that pace, limited permissions, automated checks, gradual rollouts, and fast recovery matter more, because they give early feedback and reduce the reliance on manual approvals.
The 2025 DORA research describes AI as something that amplifies both an organization’s strengths and its weaknesses. I see a clear job for SRE there: preparing the platform and the processes so the extra speed does not come with more incidents. AI itself can help too, in testing and in investigating failures.
SRE, DevOps, and Platform Engineering
The boundaries between these roles change from one organization to another, and I draw on all three. I understand DevOps as culture and collaboration between the people who develop and the people who operate software, along with engineering practices both sides share. SRE offers concrete ways to apply those principles, with measurable objectives and shared responsibility. Platform Engineering brings the product mindset to the platform.
SRE began when Ben Treynor Sloss took charge of a production team at Google in 2003 and started treating operations as a software engineering problem. Many of the ideas on this page come from the SRE books Google makes freely available, which I use as a reference and adapt to each organization.