Quick overview
This workflow runs on a schedule to health-check a web app, and when it detects an outage it gathers recent GitHub context and source files, asks Anthropic Claude to propose a fix, then creates a GitHub branch and pull request and notifies Slack via an incoming webhook.
How it works
- Runs every 2 minutes on a schedule and requests the configured health-check URL.
- Evaluates the HTTP response against the expected status code and stops if the app is healthy.
- When unhealthy, fetches recent commits and GitHub Actions workflow runs for the configured repository and branch.
- Selects a small set of candidate source files and retrieves their contents from GitHub.
- Sends the outage details, recent GitHub context, and fetched file contents to Anthropic Claude to generate a root-cause diagnosis, PR text, and file change proposals.
- If Claude proposes changes, creates a new GitHub branch, commits the updated file contents to that branch, and opens a pull request into the base branch.
- Posts an incident message to Slack with the diagnosis and PR link, or posts a separate Slack message when no confident fix is proposed.
Setup
- Create a GitHub API credential with access to read repository contents and create branches, commits, and pull requests, then set the repo owner/name and base branch in the configuration values.
- Add an HTTP Header Auth credential for the Anthropic API (and set the Anthropic model if needed).
- Replace the Slack incoming webhook URL in both Slack notification requests with your own Slack webhook endpoint.
- Update the health-check URL, expected status code, timeout, and max diagnostic file limit to match your application and desired sensitivity.