91̽ Thu, 02 Jul 2026 13:34:02 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.3 Meet SteadyBuddy: Your AI companion for reliability testing. /blog/meet-steadybuddy-your-ai-companion-for-reliability-testing/ Thu, 02 Jul 2026 10:06:21 +0000 https://steadybitstg.wpenginepowered.com/?p=2277 Building reliable systems requires more than monitoring production. You need to continuously validate how your applications behave under real-world conditions, uncover weaknesses before users do, and verify that resilience improvements actually work.

The challenge is knowing what to test, how to build meaningful experiments, and how to interpret the results.

Today, we’re excited to introduce SteadyBuddy, 91̽’s AI-powered assistant that helps you design, execute, and understand reliability tests using natural language.

Instead of navigating experiment builders or searching through previous runs, simply tell SteadyBuddy what you want to validate. It can help you create tailored reliability experiments, explain results, and recommend your next tests, all while respecting your existing permissions and keeping you in control. Want to listen, watch, and learn? Then check out this video, or read on if you prefer.

From Questions to Reliability Tests in Seconds

SteadyBuddy brings conversational AI directly into the 91̽ platform.

Want to validate how your checkout service behaves when a Kubernetes pod disappears? Just ask:

“Test how my checkout service behaves when one of its pods is killed.”

SteadyBuddy understands your request, asks clarifying questions when necessary, and generates a complete experiment draft tailored to your selected environment. From there, you can review and refine it in the experiment editor or run it immediately.

No hunting through actions. No manually assembling experiments. Just describe what you want to learn about your system.

Know What to Test Next

One of the biggest challenges in reliability engineering isn’t running tests. It’s deciding which scenarios provide the most value.

SteadyBuddy analyzes the targets available in your selected environment and recommends reliability tests that are relevant to your infrastructure. Rather than generic suggestions, you receive ready-to-run experiment ideas based on the services, workloads, and resources that actually exist in your environment.

Whether you’re onboarding a new application, validating a production environment, or expanding your reliability program, SteadyBuddy helps identify meaningful opportunities to improve resilience.

Understand Results Faster

Running a reliability test is only half the story.

After an experiment completes, SteadyBuddy can help explain what happened by analyzing the actual execution results.

Instead of digging through logs and timelines, simply ask questions like:

“Why did my last experiment fail?”

SteadyBuddy reviews the relevant experiment run, identifies what happened during execution, and explains the outcome using the available run data. The answers are grounded in the experiment’s actual results, making it easier to understand failures, identify weaknesses, and determine your next steps.

Built with Security and Control in Mind

Introducing AI into engineering workflows requires trust.

That’s why SteadyBuddy operates within the exact same security boundaries as the rest of the 91̽ platform.

SteadyBuddy only has access to the information you’re already allowed to see, including your team’s environments, available actions, reliability advice, experiment history, and the resources you have permission to access.

Just as importantly, SteadyBuddy cannot perform destructive operations on its own.

Today, it operates in a read-only capacity. It can analyze information, answer questions, and prepare experiment drafts, but every action that changes your environment requires explicit human approval. Creating, modifying, or running experiments always remains under your control.

As SteadyBuddy evolves, we’ll continue introducing additional safeguards before enabling more advanced capabilities.

Flexible Deployment for Every Environment

Whether you use 91̽ SaaS or run 91̽ on-premises, SteadyBuddy fits naturally into your deployment model.

For SaaS customers, SteadyBuddy is available as an opt-in feature that can be enabled by a tenant administrator. AI requests are processed using Anthropic models, and customer data is never used to train foundation models.

For self-hosted deployments, you stay fully in control. SteadyBuddy connects to your own configured AI provider, ensuring that data remains within your infrastructure rather than passing through any 91̽-hosted AI service.

In both deployment models, SteadyBuddy only sends the context necessary to answer your request and never transmits credentials or secrets.

AI That Understands Your Reliability Platform

Unlike general-purpose AI assistants, SteadyBuddy understands your reliability testing workflows.

It knows about your environments, available actions, target types, experiment history, and reliability advice. That context allows it to provide recommendations and explanations that are specific to your systems, instead of generic guidance that still requires significant manual work.

The result is an assistant that helps experienced reliability engineers move faster while making reliability testing more accessible to platform, DevOps, and application teams.

Available Today

SteadyBuddy is available now for eligible 91̽ customers.

Please complete this form, and we will get you enabled. Once we get you rocking and rolling, you can:

  • Generate tailored reliability test ideas for any environment.
  • Create complete experiment drafts using natural language.
  • Analyze experiment runs and understand failures faster.
  • Explore new ways to validate system resilience without manually building every experiment.

We’re excited to see how SteadyBuddy helps teams spend less time configuring experiments and more time improving system reliability.

The future of reliability testing isn’t replacing engineers. It’s giving them a faster, more intuitive way to validate resilient systems with confidence.

]]>
Introducing 91̽ Labs: Help Shape the Future of Reliability Testing /blog/introducing-steadybit-labs-help-shape-the-future-of-reliability-testing/ Thu, 02 Jul 2026 09:58:46 +0000 https://steadybitstg.wpenginepowered.com/?p=2275 Since day one, 91̽ has been driven by a vision that extends beyond building a product. We’ve always believed reliability testing should be an integral part of how modern software is built, validated, and delivered.

Over the years, we’ve helped engineering teams at some of the world’s largest organizations make reliability a continuous practice rather than a one-time exercise. But as the way we build software evolves, so does our vision for what’s next.

Today, we’re excited to introduce 91̽ Labs.

91̽ Labs is an exclusive community for customers and practitioners who want early access to the innovations shaping the future of reliability testing. It’s a place where we’ll share experimental capabilities, gather feedback from real engineering teams, and collaborate on the next generation of the 91̽ platform.

The pace of change in software engineering has never been faster. AI is transforming how teams develop, operate, and validate software, creating new opportunities to improve reliability while introducing entirely new challenges. We believe the best way to navigate this shift is by building alongside the engineers solving these problems every day.

As a member of 91̽ Labs, you’ll have the opportunity to:

  • Get early access to experimental features before they’re generally available.
  • Explore new AI-powered capabilities as they evolve.
  • Influence our product direction through direct feedback.
  • Collaborate with the 91̽ product and engineering teams.
  • Help shape the future of reliability testing.

Some ideas will graduate into core platform capabilities. Others may evolve in unexpected directions. That’s exactly what Labs is for: a space to experiment, learn, and innovate together.

If you’re excited about the future of reliability engineering and want to help define what’s next, we’d love to have you join us.

Register your interest in becoming a member of 91̽ Labs today and help shape the future of reliability testing.

]]>
Introducing Advanced Reporting in 91̽: Turn Reliability Testing into Measurable Progress /blog/introducing-advanced-reporting-in-steadybit-turn-reliability-testing-into-measurable-progress/ Wed, 10 Jun 2026 06:16:15 +0000 https://steadybitstg.wpenginepowered.com/?p=2250

 

]]>
Introducing Services in 91̽ to Enable Turnkey Reliability Testing and Risk Assessment /blog/introducing-services-in-steadybit-to-enable-turnkey-reliability-testing-and-risk-assessment/ Tue, 24 Mar 2026 21:59:38 +0000 /?p=2227 At 91̽, we enable organizations to learn about their system reliability daily through automatic risk detection and controlled chaos experiments. If you conduct this type of proactive testing, you can improve your resilience and prevent costly outages.

But in practice, making the shift to proactive testing isn’t easy.

With each new alert, reliability engineers are drawn into a series of reactive tasks, from alert triage and root cause analysis to vulnerability patching and resource optimization.

How can you roll out reliability testing quickly while staying focused on making the most impactful resilience improvements?

Today, we are launching a new feature called Services to solve this challenge.

Prioritizing Reliability Risks with Defined Services

With the ongoing shift to microservices and the advent of agentic AI, modern software teams are often operating and maintaining dozens or hundreds of services at a time.

Organizations can gauge their overall reliability posture by assessing the current level of risk that each of their services pose to business operations.

What are Services in 91̽?

We heard from customers that a service-centric view aligns better with how they approach their daily work and connect proactive reliability work to business value.

By adding in 91̽, you can now easily view test coverage, run associated reliability tests with a click, and surface the most business-critical risks.

When you install 91̽, our agents automatically discover potential targets for experiments (e.g. containers, hosts, Kubernetes deployments).

To set up services, you can use our query language or UI to define different groups of targets as services.

Defining Services in 91̽

If you already have a service catalog in a tool like , you can use the 91̽ API to automate this service creation.

You can also enrich services with additional context using custom properties. For example, it might be helpful to label services by business criticality or include links to external observability dashboards.

Services - Custom Properties

With this catalog of services in 91̽, it’s easy to see which services have sufficient testing coverage and prioritize reliability risks based on potential business impact.

What are Service Profiles?

Once you have defined services in 91̽, you can instantly pull in relevant experiments by choosing a service profile.

A consists of experiment templates organized into use case categories.

You can assign relevant experiment templates to categories like Scalability, Redundancy, and Dependencies.

When you assign a service profile to a service, you’ll see those experiments queued up in the service view, organized by category tabs. This makes it easy to rapidly expand your testing coverage and identify performance gaps faster.

Service Profiles - Experiments

With services and service profiles, you can now roll out chaos engineering across your organization with a flexible turnkey approach.

The Business Impact of Addressing Reliability Risks

With this new service-centric approach, it’ll be easier to connect the dots between chaos experiments and their impact on business operations and uptime.

ROI calculators like this one we released recently estimate the average savings that chaos engineering can deliver as you scale reliability testing coverage.

We’re excited to embed these savings assessments into 91̽ directly so you’ll be able to quantify reliability improvements as you mitigate business risks.

If you want to hear more about the services feature or get a preview of what we’re building next, just reach out to book a call with our team. You can also start a of 91̽ to explore how our platform enables chaos engineering at scale.

]]>
Using Grafana and 91̽ MCP Servers in LLM-Based Reliability Workflows /blog/using-grafana-and-steadybit-mcp-servers-in-llm-based-reliability-workflows/ Tue, 17 Mar 2026 14:17:58 +0000 https://steadybitstg.wpenginepowered.com/?p=2209 Leading observability platforms like provide engineering teams with powerful visualization and monitoring capabilities that turn aggregated logs into useful system intelligence. These comprehensive dashboards and alerting systems are essential for assessing and improving an organization’s reliability posture, but they become even more powerful when they are utilized in workflows with other tools.

With the emergence of MCP (Model Context Protocol) servers, commercial tools can now easily connect with LLM tools like Claude and Gemini to create novel workflows that leverage specific data sets and capabilities. By connecting complementary MCP servers, you can create powerful workflows that would have previously been too cost prohibitive to attempt with traditional point-to-point integrations.

While observability provides a view of the history of your system performance, chaos engineering is a proactive approach for testing performance. Instead of waiting for certain conditions to occur in production, teams can run chaos experiments to simulate an event or series of events to see how their systems react and respond. With a tool like 91̽, it’s easy to design these experiments and deploy them safely across designated targets.

Now that Grafana and 91̽ have both launched MCP servers, it’s worth exploring how these two platforms can enable teams to develop innovative combined workflows.

Combining Grafana and 91̽ Capabilities

When you connect the with , you unlock powerful new reliability tactics. AI-powered analysis can consider data and outcomes from both systems, including current health metrics, dashboard visualizations, alert configurations, and experiment results.

Grafana brings comprehensive observability visualization—customizable dashboards, alert rules, data source integrations, and historical trend analysis. 91̽ contributes chaos experiment results, system resilience insights, and proactive testing capabilities. Together, they enable teams to create a complete picture of their reliability posture with both reactive monitoring and proactive testing strategies.

To make this more tangible, we’ve created some sample reliability workflows that put both MCP Servers to work with LLM prompts.

Real-World AI-Powered Reliability Workflows

Here are some practical examples of how SRE teams could leverage both Grafana and 91̽ MCP Servers in combined workflows.

Strategic Experiment Planning

“Review all the critical incidents in Grafana in the last 90 days. Using that info, can you recommend what services we should focus on with our 91̽ experiments to improve our reliability?”

This prompt enables AI to analyze historical incident patterns and suggest targeted chaos experiments. Instead of guessing which services need resilience testing, you receive data-driven recommendations based on actual failure patterns.

SLO-Based Experiment Design

“Look at all the SLOs defined in Grafana, then use each SLO target to create a hypothesis that could be used for a chaos experiment. Export them as an excel file.”

This workflow transforms your existing Service Level Objectives into actionable chaos experiment hypotheses. The AI identifies potential failure modes that could impact each SLO and suggests specific experiments to validate your system’s resilience against these scenarios.

Reliability Impact Analysis

“Can you create a report, from our first experiment execution in 91̽ to now, that shows the change in how many incidents are occurring in Grafana per month?”

Track the effectiveness of your chaos engineering program by correlating experiment execution with incident reduction. This simplifies the reporting burden on tracking and measuring the impact of your chaos engineering efforts on system reliability.

Incident-Driven Experiment Creation

“Use the last critical incident in Grafana to create a new experiment for 91̽ to run, written in JSON, that replicates the system conditions that caused the issue.”

Transform post-incident learnings into proactive resilience testing. By analyzing the conditions and details of the most recent critical incident, an LLM can generate a specific experiment in JSON that replicates the conditions that led to the failure. With this new experiment generated, you can run it regularly as a regression test to validate that a fix has successfully been implemented and remains effective over time.

Lowering the Chaos Engineering Learning Curve

With a natural language interface, it’s easier for engineers of all skill levels to make specific queries and get meaningful results. For chaos engineering beginners, this method enables interactive learning as they are able to ask questions of the LLM directly. While it’s not a substitute for team knowledge sharing and training, it can help lower the learning curve.

Veteran engineers with plenty of chaos engineering experience can now more rapidly ask questions without the limits of query languages or complex dashboards. Ideas that would have required weeks or months of planning can be prototyped much faster. This democratizes access to reliability insights across the entire team.

Uniting chaos engineering insights from 91̽ with comprehensive observability data from Grafana leads to a powerful feedback loop, democratizing access to reliability insights across the entire team. Each experiment provides sample data for how the systems will respond under a specific stress, and then you can validate resilience improvements with real-world observability data.

Getting Started with Grafana and 91̽ Workflows

While some teams are opting for fully agentic SRE tools, MCP servers offer an approach that ensures there is a human-in-the-loop and an expert is able to learn from and make decisions with each insight. The combination of Grafana’s observability expertise and 91̽’s chaos engineering platform creates a powerful foundation for building more reliable systems and more resilient teams.

Ready to build your own LLM-powered reliability workflows?

You can get started by scheduling a demo with 91̽ to get more proactive with your approach to reliability today.

]]>
How to Prepare Your Services to Handle Availability Zone Outages /blog/how-to-prepare-your-services-to-handle-availability-zone-outages/ Tue, 02 Dec 2025 17:00:40 +0000 /?p=1882 Availability zone outages are rare but incredibly disruptive. On , a DNS issue in Amazon’s US-EAST-1 region resulted in outages across all of its 5 zones. Critical applications stopped working and customers were left stranded across industries. Once again, engineering teams saw just how vulnerable their services could be to dependencies.

Following the outage, some people took the opportunity to advocate for a multi-cloud or multi-region strategy. Others seemed to throw their hands up and accept that outages like this are just an inevitable result of any modern cloud architecture.

Regardless of your configuration, Availability zone (AZ) outages will occur. Some degradation may be unavoidable. By testing for chaotic conditions proactively, you can ensure that your services still deliver an optimal customer experience.

Reliability can be a true competitive advantage in keeping customers happy and winning new ones.

In this article, we’ll explain how you can use chaos engineering principles to run availability zone outage experiments on your systems. Don’t guess how your systems will respond. Test it.

Distributing Systems Across Availability Zones

If you run all of your services from within one availability zone, an availability zone outage will disable your entire system. That’s why it’s a reliability best practice to have coverage across multiple availability zones, so when one AZ goes down, traffic can be rerouted.

During rerouting, the load on the remaining Availability Zones (AZs) will increase. If only one AZ remains, it must absorb the full scale-up on its own. If multiple AZs remain, they can distribute the increased demand, allowing for a smoother transition. However, this works effectively only if each AZ has been provisioned with enough headroom to absorb the full projected load in the event of an AZ loss.

Paying for resources in multiple availability zones can be expensive, but outages impacting customers can be equally or much more expensive. You can read more about the benefits of using at least 3 availability zones in .

By testing how your systems perform their failover processes, you can better understand where to invest to efficiently maximize your performance and reliability.

Simulating Availability Zone (AZ) Outages with Chaos Engineering

Proactive reliability tests, or chaos experiments, intentionally inject failures in your systems in a controlled, safe way. To simulate an availability zone outage, you can run what’s called a blackhole attack.

What is a Blackhole Attack?

A blackhole attack simulates a complete network outage within a specific availability zone (AZ) or subnet. It effectively drops all incoming and outgoing network traffic for the targeted area, creating a “blackhole” where services become unreachable.

Real-world events that can mimic a blackhole zone attack include:

  • A misconfigured network access control list (ACL) that accidentally blocks all traffic.
  • A physical network hardware failure within a data center.
  • A widespread cloud provider outage affecting an entire availability zone.

By simulating this attack, you can see exactly how your system behaves when a significant part of its infrastructure suddenly goes dark.

More specifically, you can check if your traffic failover process works correctly, how the remaining services handle the increased load, and whether your monitoring tools detect the issue and raise an alert in a timely manner.

Setting Up Tooling to Run Chaos Experiments

For this walkthrough, we will use 91̽, a leading chaos engineering platform with a drag-and-drop editor for designing and running experiments quickly.

Before you begin, ensure you have the following prerequisites in place:

  • Access to 91̽: You will need an active 91̽ account, installed via SaaS or On-prem. You can try 91̽ for free for 30 days by .
  • Installing Agents & Cloud Extensions: If you want to test your environments directly, you will need to first install the 91̽ agents, one per network boundary. Next, you should install the extension for the cloud provider you use: , , or .
  • Install Monitoring Extensions: You can check for alerts and gather more thorough logs by installing the extension for your particular Observability tool: , , , , and .

You can use open source scripts to run experiments like this, but that approach makes it more challenging to deploy experiments and review results.

Designing Blackhole Zone Attack

Once you have 91̽ set up, you’re ready to build your first blackhole experiment. Here are the next steps to take:

Step 1: Choose Targets for Your Experiment

By installing the 91̽ agent and related extensions, 91̽ is able to automatically discover your cloud resources like AWS EC2 instances and subnets. You can review your targets and identify the AZ or subnet(s) you want to isolate.

We generally recommend starting your experiments with non-production environments (e.g. “Dev” or “QA”) to avoid negatively impacting real users.

Step 2: Design the Experiment

In the 91̽ UI, navigate to the Experiment Editor to design your experiment. You can either use a or build from scratch. If you want to start with a blank canvas, here’s what you would do:

  1. Define the Attack: Add a new network attack by selecting the action and dragging it into the canvas.
  2. Select the Target: Choose which AWS Availability Zones you want to target with this attack. 91̽ will show you all discovered targets with related metadata like region or AZ, which helps make selecting targets easy.
  3. Set the Duration: The experiment will block traffic by creating and assigning a new network ACL to every subnet in every VPC in the selected zone. For your first run, a short duration like 60 seconds is often sufficient to observe the initial impact without causing prolonged disruption.
  4. Add Checks: You can add checks to run over the course of your experiment, like a periodic HTTP request on a certain endpoint. You could also add a check to see whether your monitoring tool raises relevant alerts as expected.

You can see an example of this type of experiment here. The experiment is designed to check that the HTTP request is successful initially before start the blackhole zone attack. There is a wait step to account for recovery time. Then, the HTTP check resumes to ensure that the service has recovered in the amount of time expected:

Experiment Template:

If you want to run an experiment that only simulates a subnet outage, you would take the same steps, except you would select the action instead and specify an AZ to take out during the experiment.

Step 3: Execute the Experiment

When you are happy with your experiment design, you can hit “Run Experiment” to watch it run in real-time.

As the experiment executes, 91̽ will apply the blackhole by modifying the network ACLs to deny all inbound and outbound traffic for the selected zone(s). Throughout the experiment, any checks that you set up will allow you to actively monitor your system’s key performance indicators (KPIs). Referencing your observability dashboard can also be helpful for things like error rates, latency, and resource utilization in the remaining active zones.

Step 4: Review the Performance Results

Once the experiment is finished, 91̽ will automatically revert all changes, restoring network traffic to the targeted subnet. Now, you can analyze exactly how your systems handled the outage.

  • Did your application traffic successfully failover to the healthy availability zones?
  • Did your monitoring systems trigger the expected alerts?
  • Did any services experience degraded performance or crash entirely?
  • Were there any unexpected cascading failures?

For example, if you found that traffic did not failover, you might need to adjust your load balancer or Kubernetes service configurations. If alerts didn’t fire, you may need to make updates in your observability tool.

Whatever the results are, you now have real data you can review to see if you can make any improvements.

Step 5: Continuous Verification

After implementing any fixes, you should run the experiment again to verify that your changes have resolved the problems that you intended to solve.

Since systems are constantly changing, this iterative cycle of testing, learning, and improving is critical to building truly resilient systems. To accomplish this, some teams build experiments into their CI/CD workflows as a quality gate to maintain reliability standards over time.

By identifying and validating the limits of your systems, you can map out key break points and know how your services will react to different conditions.

Adopting Proactive Reliability Practices

By proactively testing for failures, you can ensure your services remain highly available and delivering for your customers. Failures are inevitable, but being prepared for failures is a choice that your organization actively makes.

Are you ready to test the resilience of your systems?

Get started with a today to explore how chaos experiments can upgrade the reliability of services across your organization.

]]>
Unlocking New Reliability Workflows with the Datadog and 91̽ MCP Servers /blog/unlocking-new-reliability-workflows-with-the-datadog-and-steadybit-mcp-servers/ Wed, 15 Oct 2025 19:26:53 +0000 https://steadybitstg.wpenginepowered.com/?p=1815 Leading observability tools like provide engineering teams with one central plane for reading and visualizing system performance. These aggregated logs and monitoring data are essential for assessing and improving an organization’s reliability posture, but they become even more powerful when they are utilized in workflows with other tools.

With the emergence of MCP (Model Context Protocol) servers, commercial tools can now easily connect with LLM tools like Claude and Gemini to create novel workflows that leverage specific data sets and capabilities. By connecting complementary MCP servers, you can create powerful workflows that would have previously been too cost prohibitive to attempt with traditional point-to-point integrations.

While observability provides a view of the history of your system performance, chaos engineering is a proactive approach for testing performance. Instead of waiting for certain conditions to occur in production, teams can run chaos experiments to simulate an event or series of events to see how their systems react and respond. With a tool like 91̽, it’s easy to design these experiments and deploy them safely across designated targets.

Now that Datadog and 91̽ have both launched MCP servers, it’s worth exploring how these two platforms can enable teams to develop innovative combined workflows.

Utilizing Datadog and 91̽ in LLM Workflows

When you connect the Datadog MCP server with 91̽’s MCP server, you unlock powerful new reliability tactics. AI-powered analysis can consider data and outcomes from both systems, including current health metrics, SLOs, incident alerts, and the experiment results.

Datadog brings comprehensive observability data—metrics, logs, traces, and incident history. 91̽ contributes chaos experiment results, system resilience insights, and proactive testing capabilities. Together, they enable teams to create a complete picture of their reliability posture.

To make this more tangible, we’ve created some sample reliability workflows that put both MCP Servers to work with an LLM prompt.

Real-World AI-Powered Reliability Workflows

Here are some practical examples of how SRE teams could leverage both Datadog and 91̽ MCP Servers in combined workflows.

Strategic Experiment Planning
“Review all the critical incidents in Datadog in the last 60 days. Using that info, can you recommend what services we should focus on with our 91̽ experiments to improve our reliability?”

This prompt enables AI to analyze historical incident patterns and suggest targeted chaos experiments. Instead of guessing which services need resilience testing, you receive data-driven recommendations based on actual failure patterns.

SLO-Based Experiment Design
“Look at all the monitor-based SLOs defined in Datadog, then use each SLO to create a hypothesis that could be used for a chaos experiment, and export them as a list.”

This workflow transforms your existing Service Level Objectives into actionable chaos experiment hypotheses. The AI identifies potential failure modes that could impact each SLO and suggests specific experiments to validate your system’s resilience against these scenarios.

Reliability Impact Analysis
“Can you create a report, from our first experiment execution in 91̽ to now, that shows the change in how many incidents are occurring in Datadog per month?”

Track the effectiveness of your chaos engineering program by correlating experiment execution with incident reduction. This simplifies the reporting burden on tracking and measuring the impact of your chaos engineering efforts on system reliability.

Incident-Driven Experiment Creation
“Use the last critical incident in Datadog to create a new experiment for 91̽ to run, written in JSON, that replicates the system conditions that caused the issue.”

Transform post-incident learnings into proactive resilience testing. By analyzing the conditions and details of the most recent critical incident, an LLM can generate a specific experiment in JSON that replicates the conditions that led to the failure. With this new experiment generated, you can run it regularly as a regression test to validate that a fix has successfully been implemented and remains effective over time.

Lowering the Chaos Engineering Learning Curve

With a natural language interface, it’s easier for engineers of all skill levels to make specific queries and get meaningful results. For chaos engineering beginners, this method enables interactive learning as they are able to ask questions of the LLM directly. While it’s not a substitute for team knowledge sharing and training, it can help lower the learning curve.

Veteran engineers with plenty of chaos engineering experience can now more rapidly ask questions without the limits of query languages or complex dashboards. Ideas that would have required weeks or months of planning can be prototyped much faster. This democratizes access to reliability insights across the entire team.

Uniting chaos engineering insights from 91̽ with comprehensive observability data from Datadog leads to a powerful feedback loop, democratizing access to reliability insights across the entire team. Each experiment provides sample data for how the systems will respond under a specific stress, and then you can validate resilience improvements with real-world observability data.

Getting Started with Datadog and 91̽ Workflows

While some teams are opting for fully agentic SRE tools, MCP servers offer an approach that ensures there is a human-in-the-loop and an expert is able to learn from and make decisions with each insight. The combination of Datadog’s observability expertise and 91̽’s chaos engineering platform creates a powerful foundation for building more reliable systems and more resilient teams.

Ready to build your own LLM-powered reliability workflows?

You can get started by scheduling a demo with 91̽ to see what a proactive reliability program can look like for your organization.

]]>
How to Check Kafka Consumer’s Reaction to Record Loss /blog/how-to-check-kafka-consumers-reaction-to-record-loss/ Wed, 15 Oct 2025 19:13:03 +0000 https://steadybitstg.wpenginepowered.com/?p=1812 Apache Kafka is a key technology in modern data streaming, enabling systems to handle and transit massive volumes of data. Kafka consumers are applications that subscribe to topics in an Apache Kafka cluster to read and process data records. As endpoints of Kafka’s publish-subscribe model, they retrieve messages published by producers.

But what happens when things go wrong?

For Kafka consumers, record loss can be a critical failure point, leading to data inconsistencies and significant business impact. The only way to ensure your systems are resilient is to proactively test how your consumers behave when faced with this exact scenario.

In this guide, we will walk through how to design and run an experiment to check your Kafka consumer’s reaction to record loss. We’ll show you how to simulate this failure safely, monitor the results, and gain confidence in your system’s reliability using chaos engineering.

What is Record Loss in Kafka?

Record loss occurs when a consumer fails to process messages that have been produced to a topic. This isn’t just a theoretical problem; it can happen in several ways in a production environment.

The most common causes include:

  • Authorization Issues: A consumer might suddenly lose permission to access a specific topic, preventing it from fetching new records.
  • Offset Mismanagement: Incorrect handling of consumer offsets can cause the consumer to “jump” over a batch of records, effectively losing them.
  • Intentional Deletions: An administrator might delete records from a partition, for example, to comply with data retention policies, while a consumer is temporarily offline.

The impact of even a small amount of record loss can be severe. It can corrupt downstream data, break application logic, and lead to poor user experiences. That’s why verifying your consumer’s resilience is not just a best practice—it’s essential for operational readiness.

Testing Kafka Resilience with Chaos Engineering

To effectively test your Kafka consumer, you need to run meaningful tests on your system. Instead of waiting for issues to occur in Production, you can run experiments to proactively observe how your systems respond to record loss issues.

For this guide, we’ll build our experiment using 91̽, the chaos engineering platform that makes it easy to reveal reliability gaps and design experiments fast.

Before you begin, you will need:

  • An active Kafka cluster and a consumer application.
  • Access to the 91̽ platform.

You can use the open source Kafka extension to connect 91̽ with your Kafka cluster and start running valuable experiments in minutes.

Designing a Kafka Experiment to Simulate Record Loss

Our experiment should realistically mimic the conditions that lead to record loss. The goal is to see how the consumer application responds when records disappear from a topic partition between fetches.

You could start with hypothesis like:

“Simulate the loss of records to see how the consumer recovers. We should look at the logs of the consumer after the deny to see that it begins to consume records with value corresponding only to the last records after the delete.”

Then, these are the key actions to orchestrate this use case:

  1. Deny Topic Access: We’ll start by temporarily blocking the consumer’s access to the topic. This simulates a permission issue and pauses consumption.
  2. Produce Messages: While the consumer is blocked, we will continue to produce messages to the topic to simulate a live environment.
  3. Delete Records & Adjust Offset: This is the core of the attack. We will delete the newly produced records and manipulate the partition’s offset to make it seem as if the records never existed.
  4. Restore Access: Lastly, we’ll restore the consumer’s access and observe its behavior.

This sequence creates a scenario where the consumer, upon reconnecting, is faced with a gap in the message log. By seeing how it handles this gap, you can identify potential performance issues.

This pre-built experiment template from 91̽’s Reliability Hub shows what that would look like:

91̽ Experiment Template:

This experiment creates a scenario where the consumer, upon reconnecting, is faced with a gap in the message log. By running it, we can see how the consumer handles this gap.

Reviewing the Results and Consumer Reactions

With the experiment running, your focus shifts to observation. In a chaos engineering tool like 91̽, you’ll be able to watch as each step executes and see how your application responds.

When it is complete, check your application’s logs for any errors or warnings related to authorization failures or offset problems. You can also monitor the consumer lag for your topic. This metric shows the difference between the latest offset on the topic partition and the offset your consumer has committed. A sudden drop or unusual fluctuation in lag can indicate that the consumer has skipped records.

If the consumer logs show that after the deny, the consumer began to consume records with values corresponding only to the last records after the delete, then your experiment was a success and you have a resilient consumer.

A resilient consumer detects the offset jump, logs a critical error, and either shuts down safely or triggers an alert for manual intervention. It would not silently continue processing from the new offset, which would hide the data loss.

If your experiment reveals that you have a vulnerable consumer, you have a clear action item: improve your consumer’s error-handling logic. This could involve adding checks for offset continuity or implementing more robust monitoring and alerting around consumer lag.

Getting Started Running Kafka Reliability Tests

Waiting for a production incident to discover a weakness in your Kafka consumers is risky and reactive. By proactively testing for scenarios like record loss, you can identify and fix issues before they impact your customers. Chaos engineering is the best practice for building this type of operational resilience.

Ready to find out how your Kafka consumers stand up to pressure?

You can get started with a of 91̽ or book a demo for your team today.

]]>
How to Test Kubernetes Deployment Degradation When RabbitMQ Is Down /blog/how-to-test-kubernetes-deployment-degradation-when-rabbitmq-is-down/ Thu, 18 Sep 2025 18:23:19 +0000 /?p=1774 Do you know how your application will perform if your RabbitMQ cluster goes down?

In this post, we will outline how you can proactively test your systems to validate graceful degradation when RabbitMQ is unavailable. You will learn why this is essential for system reliability to conduct testing like this and how to run a chaos experiment to ensure your systems can handle this type of failure.

What is RabbitMQ and Why is Graceful Degradation Essential?

Many teams use , an open-source messaging and streaming broker, to coordinate the data communication for their applications. It’s a vital message broker in many architectures, especially within Kubernetes environments where it facilitates asynchronous communication between microservices.

The Role of RabbitMQ in Kubernetes

In a Kubernetes-native architecture, microservices need a way to communicate without being tightly coupled. RabbitMQ functions as a post office for your services. It manages workloads, enables scalable asynchronous operations, and ensures messages are delivered reliably.

Common use cases include:

  • Background Job Processing: Offloading long-running tasks like generating reports or sending emails, so the main application remains responsive.
  • Event-Driven Architectures: Allowing services to react to events happening in other parts of the system without direct dependencies.
  • Inter-Service Communication: Managing the flow of data between dozens or even hundreds of microservices.

Defining Graceful Degradation

Graceful degradation is a system’s ability to maintain limited but essential functionality when a critical dependency fails. Instead of a catastrophic failure that brings everything down, the system sheds non-essential features while keeping core services online.

When a RabbitMQ cluster becomes unavailable, the application should continue to operate with reduced capability, protecting the user experience from a complete outage, and the outage should be detected and flagged by your observability tool.

Consider an e-commerce application. If RabbitMQ, which processes new orders for the shipping department, goes down, graceful degradation means customers may still browse products and add items to their cart. They just might see a temporary message that order processing is delayed, but they are not met with an error page.

The Importance of Proactive Validation

How do you know if your Kubernetes deployment will degrade gracefully? Without testing it out, you are relying on hope. By running an experiment proactively, you can safely test how your system responds and make improvements before an outage occurs.

This type of validation is critical for several reasons:

  • System Reliability: Verifying resilience prevents minor issues from escalating into major incidents that breach Service Level Objectives (SLOs).
  • User Experience: It protects users from disruptive outages, maintaining their trust in your service.
  • Operational Readiness: It prepares your team for real-world failures, reducing Mean Time To Recovery (MTTR) and the stress of firefighting incidents in production.

Now, let’s explore how to create and run an experiment for this use case.

How to Run an Experiment to Test RabbitMQ Failure

The most effective way to validate graceful degradation is with chaos engineering. By intentionally injecting a controlled failure, you can observe how your system behaves and identify weaknesses before they impact users. A chaos experiment provides the proof you need to build true system resilience.

Here’s how you can design and execute an experiment to simulate a RabbitMQ failure.

Step 1: Define Your Hypothesis and Steady State

Every good experiment starts with a clear hypothesis. This is a statement about what you expect to happen. Before you start, you must also understand your system’s “steady state,” or its normal behavior.

  • Hypothesis: “When the RabbitMQ service is unavailable, the primary application will remain operational, and user-facing services that do not depend on messaging will continue to function without error. Services dependent on RabbitMQ will fail predictably without causing a cascading failure.”
  • Steady State: Identify Key Performance Indicators (KPIs) to monitor, such as application response time, error rates (HTTP 5xx), and resource utilization (CPU/memory). This baseline is essential for measuring the impact of the experiment.

Step 2: Design the Chaos Experiment

Next, design the experiment to test your hypothesis. The goal is to simulate a RabbitMQ outage in a way that is realistic and contained.

  • Experiment: Block all network traffic to and from the RabbitMQ pods within your Kubernetes cluster. This accurately mimics a network partition or a complete service failure.
  • Blast Radius: Limit the experiment’s scope to a specific namespace or a subset of services. This minimizes potential impact, especially when running tests in a production environment. For example, you could target only the services (or subset of services) in the dev or staging namespace.

Here’s an example of the steps for this type of experiment:

91̽ Experiment Template:

You can find this experiment template in the , 91̽’s open source library of chaos engineering components, including over 80 similar experiment templates.

Step 3: Execute and Monitor the Chaos Experiment

With your plan in place, it’s time to run the experiment. If you’re using a tool like 91̽, you’ll be able to watch the experiment execute in real-time.

As the experiment runs, here are some things to keep in mind:

  • Monitor application logs: Look for connection errors, retries, or unexpected exceptions. Are services entering a crash loop trying to reconnect?
  • Check non-dependent features: Verify that core functionalities unrelated to messaging are still working as expected. Can users log in? Can they navigate the site?
  • Watch your dashboards: Keep an eye on your monitoring tools. Are error rates spiking? Is latency increasing for unrelated services?

Step 4: Analyze the Results and Make Improvements

Once the experiment is complete, analyze the outcome. Did the system behave as you hypothesized?

  • Did it degrade gracefully? If yes, congratulations! You have validated a key aspect of your system’s resilience.
  • Did it fail unexpectedly? This is also a valuable outcome. Perhaps a health check on an unrelated service failed because it had an implicit dependency on RabbitMQ. Or maybe a service went into a crash loop, consuming excessive resources and threatening the stability of its node.

Based on these findings, you can take action. Remediation might involve implementing circuit breakers to stop services from endlessly retrying failed connections, adding fallback mechanisms, or improving error handling to fail more gracefully.

Building Resilience and Operational Readiness

Running chaos experiments to validate graceful degradation is not just a technical exercise; it delivers tangible business value and strengthens your team’s skillset.

With one proactive test, you can find and fix issues in a controlled environment. With experiments run regularly or integrated into a CI/CD pipeline, you can systematically eliminate potential causes of outages. This practice is fundamental to meeting and exceeding aggressive uptime targets like 99.99%.

If you want to get started running chaos experiments with a tool that makes it easy, schedule a demo for your team or . Build more resilient systems one experiment at a time.

]]>
How to Test Load Balancer Failover During an AWS Availability Zone Outage /blog/how-to-test-load-balancer-failover-during-an-aws-availability-zone-outage/ Tue, 16 Sep 2025 17:56:49 +0000 /?p=1768 If you are running applications on cloud infrastructure, an issue with your cloud provider could result in a costly outage. Are your systems ready for if an entire AWS Availability Zone (AZ) goes down?

Outage events like this often result in system downtime, distressed customer calls, and business operations grinding to a halt. But for some prepared teams, these incidents can be proof that their failover processes are working successfully.

Correctly configured load balancers could make a big difference. By setting up load balancers to respond automatically to AZ outages, you can trust that traffic will be rerouted to healthy instances in other zones without a single user noticing.

It takes both intentional failover design and continuous verification to achieve this level of system resilience. Just setting it up and assuming that it works correctly is a hope-based strategy.

In this post, we’ll explain how you can proactively test your load balancer’s cross-zone failover capabilities with chaos engineering to ensure true system resilience.

The Role of Load Balancers in Achieving High Availability

An Availability Zone (AZ) is a distinct cloud region, supported by separate, physical data centers that distribute the geological risk of any one location. This decentralized design is the foundation of high-availability architectures and common across cloud providers. In this post, we’ll focus on AWS specifically.

Load balancers, like AWS Application Load Balancer (ALB) or Network Load Balancer (NLB), sit in front of your services and distribute incoming traffic across multiple targets, such as EC2 instances or containers. In a well-architected system, these targets are spread across multiple AZs so there are still available resources available if any one AZ goes down.

In the scenario of an AZ outage, you would hope that your load balancer would run the following steps:

  1. Health Checks: The load balancer continuously sends health check requests to each registered target.
  2. Failure Detection: If any target fails to respond correctly, the load balancer marks it as unhealthy.
  3. Rerouting Traffic: When all targets within an entire AZ become unhealthy, as they would in an outage, the load balancer’s failover mechanism activates. It then stops sending traffic to the failed AZ and seamlessly reroutes traffic to the healthy targets in the remaining AZs. This ensures that even if your application has degraded performance, it still remains available and functioning for end users.

Why It’s Critical to Test Load Balancer Failover Processes

On paper, the failover process seems straightforward. But in complex, dynamic production environments, you can’t rely on the assumption or hope that systems will always respond as designed. Here’s why continuous verification is critical.

  • Configuration Drift: Your architecture diagram might be perfect, but the reality often doesn’t line up perfectly. Misconfigurations in networking rules, security groups, or health check parameters can silently accumulate over time. These subtle changes can prevent the load balancer from correctly detecting failures or routing traffic, causing the failover to fail completely.
  • Capacity Planning Issues: Let’s say the failover works, and traffic is redirected. Do the remaining AZs have enough capacity to handle the sudden influx? Without testing, you’re just guessing. A successful failover can still lead to performance degradation, increased latency, or a complete overload of the remaining instances if your auto-scaling policies aren’t tuned correctly.
  • Dependencies and “Unknown Unknowns”: Complex systems hide complex dependencies. A failover might expose a critical service, like a database replica or a caching layer, that was only running in the failed AZ. This can trigger a cascade of failures that brings down your entire application, even though the initial load balancing worked.
  • Impact on SLOs: A failed or even a slow failover directly impacts your uptime and availability Service Level Objectives (SLOs). This erodes user trust, harms your brand’s reputation, and can lead to costly violations of Service Level Agreements (SLAs).

Simulating an AWS AZ Outage with Chaos Engineering

The only way to know how your system will react to an AZ outage is to simulate one in a controlled manner. By using a chaos engineering solution like 91̽, you can simulate this type of scenario safely.

Here’s how to build an experiment that validates your load balancer failover process during an AWS Availability Zone outage.

Step 1: Design the Chaos Experiment

A well-designed experiment starts with a clear hypothesis. This statement defines what you expect to happen during and as a result of the experiment.

  • Hypothesis: “When we simulate an AWS AZ outage by blocking all network traffic to one zone, our load balancer will successfully detect the unhealthy targets and redirect all traffic to the healthy zones. Key business metrics, like error rate and latency, will remain stable.”

Next, you must precisely define the blast radius. You want to contain the failure to a specific, narrow set of targets to run a safe and meaningful experiment.

91̽’s query-based targeting makes this simple. You can select all targets within a specific AWS availability zone (e.g., us-east-1a) without needing to manually list hosts or containers.

Step 2: Simulate the AZ Outage

With your targets defined, you can configure the attack. To perfectly simulate an unreachable AZ, we use a Blackhole zone attack.

The Blackhole attack drops all network traffic to and from the targeted instances. This mimics the real-world scenario where an entire zone becomes inaccessible due to a network partition or other major failure. You can configure this Blackhole action to run for a certain duration while also running HTTP checks in parallel to gauge application functionality.

Here’s an example of this type of experiment design.

91̽ Experiment Template:

Before you run the experiment, you should also connect your chaos engineering tool with your observability tool so you can validate that alerts are being raised in real-time as you would expect.

Step 3: Run and Analyze the Results

Now, it’s time to run the experiment and observe your system’s behavior in real time. Watch your dashboards closely.

  • Did the load balancer’s health checks correctly identify and mark the targets in the affected AZ as unhealthy?
  • Was traffic successfully and quickly routed to the instances in the other AZs?
  • Did your application’s error rate, throughput, and latency remain within acceptable limits?
  • Did any downstream services experience issues due to hidden dependencies?

You can use these insights to improve and harden your system. You might discover a misconfigured health check, an under-provisioned auto-scaling group, or a critical dependency you never knew existed.

If your hypothesis was accurate, your experiment is a “success”. If your hypothesis was wrong, your experiment “failed” and you will need to review the results to either revise your hypothesis or improve your systems. You might discover a misconfigured health check, an under-provisioned auto-scaling group, or a critical dependency you never knew existed.

By finding and fixing reliability issues in a controlled experiment earlier in the software development lifecycle, you will be preventing future outages one experiment at a time.

Turning Chaos Experiments into Continuous Tests

Load balancers are fundamental to building highly available systems, but they are not a “set it and forget it” solution. Assuming that your failover mechanisms will work as expected is a risk your business can’t afford. You need to continuously test these types of processes to validate that your systems are truly resilient.

By running experiments on a schedule or with CI/CD automation, you can move from a reactive hope-based strategy to a proactive data-driven process that builds a culture of reliability. You’ll gain evidence that your systems can withstand one of the most common—and potentially damaging—failure modes in the cloud.

Ready to start running experiments? See how easy 91̽ makes it to create and run experiments that provide real insights on your system’s resilience. Schedule a demo or today to put your load balancers to the test.

]]>