Video summary
DevOops - Антон Егорушков - Reliablility Map и куда она ведет
Main summary
Key takeaways
Overview
The video is a discussion in a DevOops / DevOps conference context about Reliability Map—a structured way to evaluate a company/service’s maturity in reliability engineering and plan next steps toward higher fault tolerance and availability (often framed via the “nines” model).
The speaker argues that the map is a useful framework to help teams determine:
- Where we are now
- What to do next
It also acknowledges the risk of over-engineering.
Reliability Map: purpose and structure
The Reliability Map is presented as a table/grid defined by two axes:
- Eras / maturity stages of reliability development
- Starting simple and progressively adding practices.
- Depth of engagement with streams/tools
- Some streams may be irrelevant early on, while others become necessary later.
It is explicitly connected to SRE (Site Reliability Engineering) as practiced in the real world, with reliability.engineering referenced as the primary public source.
The speaker notes that, while many talks discuss SRE/reliability, few Russian-language materials use the Reliability Map itself as a methodological framework.
Maturity “eras” (high-level walkthrough)
Demo era / early stage
- An MVP exists
- Accessibility may “seem” good, but is not measured
- Reliability goals (e.g., reaching specific “nines”) are not yet operationally validated
Progression toward advanced reliability
- Each stage introduces new practices and tooling
- The emphasis is on systematically growing capabilities, rather than jumping directly to advanced tooling.
Key concept: Customerability Engineering (CE)
A major thread is Customerability Engineering (CE), described as:
“Sustainable engineering for the customer/client.”
It is promoted by vendors (as mentioned in the discussion), such as Yandex Premium support, where the provider monitors and improves the customer’s fault tolerance across the customer’s systems, infrastructure, and services.
Mapping CE to reliability stages
- Proactive era
- Focuses on handling high loads
- DR/DRP (disaster recovery / disaster recovery planning)
- Scalability preparedness
- Jet age / development stage mentioned
- Includes checks and validation
- The speaker specifically mentions checking during pre-production and improving stability
Additional discussion points
- On-premises large deployments and enterprise software (examples mentioned included large installs such as Yandex 360, Yandex DataLens, and Bitrix-type systems)
- CE practices used to help assess/install/operate reliability patterns operationally
- Observability examples, such as Grafana dashboards, to evaluate fault tolerance/reliability
- AI-assisted checks mentioned as “AI that checks code” during pre-production to improve stability
- Acknowledgement that Yandex Cloud can still have incidents, i.e., practices are not “magical,” but still improve resilience
Cloud vs on-prem applicability
The speaker claims the practices are applicable to both cloud and on-prem.
However, the “prism” of CE tends to be more cloud-oriented (customer-cloud operations), though it can be adapted for on-prem scenarios.
Where the map “leads” (and limits)
The end target is expected to align with an “autonomous era” (high maturity), but the speaker says:
- Very few teams reach it
- Multiple required practices are still missing or not sufficiently established in real organizational practice
Caution: avoiding “practices for the sake of practices”
- Teams sometimes adopt practices for the sake of the practice itself
- This can lead to overkill (e.g., paying for complex resilience when simpler fixes and fault containment would suffice)
The map should therefore help avoid unnecessary cost and complexity.
Training/report format and practical expectations
The speaker intends a report, not a workshop/master class, because:
- Case interactivity could make stories less applicable to listeners
The expected value of the talk:
- High-level practices (explicitly framed as “least covered”)
- Clear takeaways on how to move between eras
- The mapping is treated as a framework, not just a checklist
Example content mentioned
Likely included examples (some wording is unclear due to auto-captions), such as:
- “Custumerity survey” / customer-related survey (exact phrasing unclear)
- Predictive monitoring
- Observability integration across tools (including a “purple quadrant” mentioned in the map)
- AI-assisted checks during pre-production
- Recognition that cloud providers still experience incidents, so improvements are practical—not absolute
Translation plans
The speaker discusses translating the Reliability Map content into Russian as an artifact/supplement:
- Make it readable and usable without forcing the audience to work through the full English map
- Not necessarily required as part of the main report initially
- Planned for later (potentially around September)
Naming/translation strategy
- Avoid awkward translation of abbreviations
- Keep English names where appropriate, similar to how some sources handle things like “technological radar / SWK”
- Add explanations when needed
GPT/ChatGPT discussion (tool assistance, not replacement)
A question is raised about whether chat-based AI can help with the Reliability Map.
The speaker’s view:
- GPT can help generate/interpret practices, but must be strictly controlled
- GPT may misinterpret eras and reliability-stage logic
- GPT can assist with questions like:
- “Given practice X, what should we do?”
- Suggestions for monitoring implementations
- Inferring what’s already configured
- But GPT won’t reliably define the next stage end-to-end without human guidance
Community and analogues
The Reliability Map is described as the most “alive” and community-supported:
- Actively developed on GitHub
- Weekly/community meetings mentioned: “Tuesdays”, held at Google (details partially unclear)
The speaker notes there are other maps (DevOps-related), but emphasizes that this one is the flagship for SRE/fault tolerance.
Community demand may also be low because the map isn’t needed day-to-day—it’s mainly used for planning and moving maturity forward.
Target audience outcomes
The talk is meant to help listeners:
- Locate themselves on the map (current era/position)
- Identify next steps (what to implement next, and in what order)
- Understand how changes affect the business, including the cost vs benefit discussion
Main speakers / sources (as mentioned)
- Vladimir (conference/program committee speaker opening the discussion)
- Anton Egorushkov (implied as the primary talk author; referenced multiple times)
- SRE community / reliability.engineering (key reference source)
- Google (mentioned in connection with SRE community meetings and CE concept origin)
- Yandex / VK / cloud providers (mentioned as examples for premium support and CE-like offerings)