On-Call Hero 2026: AI SRE Agent for Incident Management
Discover On-Call Hero, an AI SRE agent built for the Hacktoberfest Weekend Challenge using Temporal to ensure durable workflows during incidents.

Stock photo for illustration only, not from the actual event
- On-Call Hero is an AI SRE agent designed to investigate and respond to infrastructure incidents.
- It utilizes Temporal to maintain durable workflows, resuming seamlessly even if workers crash.
- Features include human-in-the-loop approvals, prompt-injection protection, and automatic rollbacks.
- Built using Python, Temporal, FastAPI, SSE, LiteLLM, and supports local models via Ollama.
Armansiddiqui9 has created a project named On-Call Hero as a submission for the Hacktoberfest Weekend Challenge under the theme Build for a Friend. The project is engineered to act as an AI SRE agent that assists engineering teams in investigating and reacting to critical system incidents.
The primary challenge the developer aimed to solve is reliability. When an AI agent is actively handling an incident, unexpected worker crashes or server restarts can derail standard automation pipelines. On-Call Hero addresses this vulnerability by integrating Temporal to manage system execution.

Stock photo for illustration only, not from the actual event
By leveraging Temporal, the workflow becomes durable. The agent can investigate an issue, propose a fix, wait for human authorization, execute the solution, verify the outcome, and roll back changes when necessary. Even if the worker process is abruptly terminated mid-incident, the workflow resumes precisely where it left off.
Implementing durable workflow engines like Temporal for AI agents is crucial for production-grade reliability. Since autonomous agents frequently encounter API timeouts, rate limits, or infrastructure hiccups, state persistence ensures that execution states are safely preserved without restarting diagnostic procedures from scratch.
The project's technology stack relies heavily on open-source tools, incorporating Python, Temporal, FastAPI, SSE, and LiteLLM. It also offers flexibility by running local models through Ollama, making local experimentation with durable AI agents accessible and transparent.
Security and governance have also been integrated into the system. The platform incorporates human approval gates before executing actions, alongside plan validation, prompt-injection protection, idempotency checks, budget controls, audit histories, and rollback handling.
Users can test the system locally by spinning up the local server and trying the S1 Bad deploy scenario. The complete source code is available on GitHub, and community feedback is welcomed as the developer continues to refine the project.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment