# Overview
Source: https://docs-sre.lightrun.com/connectors/collaboration
Use AI SRE from Slack and Microsoft Teams
Run investigations where your team already works. AI SRE is available in Slack and Microsoft Teams so you can ask questions, get evidence-based analysis, and share findings without leaving the conversation.
### Slack
Use **@Lightrun AI SRE** in Slack to investigate incidents from any channel. The same AI runs as in the web app—gathering evidence, correlating data, and returning root cause analysis and suggested actions directly in the channel.
* **@-mention to investigate** — Ask questions or describe incidents by mentioning **@Lightrun AI SRE** in a channel; responses stream into the channel with full markdown.
* **Same evidence and confidence** — Get the same evidence chains, confidence levels, and next-step recommendations as in the web UI.
* **Channel context** — Follow up in the same channel; the app uses the channel history for continuity.
For how to set up Slack, daily use, and troubleshooting, see [Slack](/connectors/slack-integration).
### Microsoft Teams
Microsoft Teams support is coming soon. You can request access from the **Connectors** page in the web app.
## Next steps
# Incidents & Alerting
Source: https://docs-sre.lightrun.com/connectors/incidents-alerting
Connect incident management platforms to AI SRE
Connect AI SRE to your incident management and alerting platforms to receive alerts, access incident data, and correlate with investigations.
The following connectors are available as beta features. You can request access to connectors based on your stack from the connectors screen in the web UI.
### PagerDuty
The PagerDuty integration enables AI SRE to:
* **Access incidents** — Query incident data, status, and updates
* **Correlate alerts** — Link related alerts and incidents together
* **Check on-call** — Access on-call schedules and escalation information
* **Review timeline** — Analyze incident timelines and response data
### Sentry
The Sentry integration enables AI SRE to:
* **Track errors** — Query error tracking data and issue details
* **Correlate issues** — Link related errors and issues together
* **Monitor releases** — Access release tracking and deployment data
* **Analyze performance** — Query performance monitoring data and metrics
## Next steps
# Infrastructure
Source: https://docs-sre.lightrun.com/connectors/infrastructure
Connect infrastructure platforms to AI SRE
Connect AI SRE to your infrastructure platforms to query cloud resources, cluster information, and infrastructure metrics during incident investigations.
The following connectors are available as beta features. You can request access to connectors based on your stack from the connectors screen in the web UI.
### AWS (Amazon Web Services)
The AWS integration enables AI SRE to:
* **Query EC2 instances** — Access instance status, health, and configuration
* **Analyze CloudWatch** — Query metrics and logs from CloudWatch
* **Inspect clusters** — Access ECS/EKS cluster data and service status
* **Monitor Lambda** — Query Lambda function metrics and execution data
* **Check databases** — Access RDS instance information and metrics
### Kubernetes
The Kubernetes integration enables AI SRE to:
* **Check pod status** — Query pod and container health and status
* **Monitor resources** — Access cluster resource metrics and capacity
* **Review deployments** — Query deployment and service information
* **Assess nodes** — Check node health, capacity, and availability
* **Analyze events** — Review cluster event logs and changes
### Google Cloud Platform (GCP)
The GCP integration enables AI SRE to:
* **Query compute instances** — Access Compute Engine instance status and metrics
* **Monitor services** — Query Cloud Monitoring metrics and dashboards
* **Search logs** — Access Cloud Logging data and queries
* **Inspect GKE clusters** — Query GKE cluster information and health
* **Check databases** — Access Cloud SQL instance metrics and status
## Next steps
# Knowledge Base
Source: https://docs-sre.lightrun.com/connectors/knowledgebase
Connect knowledge management systems to AI SRE
Connect AI SRE to your knowledge management systems to search documentation, runbooks, and tribal knowledge during incident investigations.
The following connectors are available as beta features. You can request access to connectors based on your stack from the connectors screen in the web UI.
### Confluence
The Confluence integration enables AI SRE to:
* **Search documentation** — Query and retrieve relevant documentation pages
* **Access runbooks** — Find and retrieve runbooks for incident response
* **Query knowledge base** — Search your knowledge base for relevant information
* **Read content** — Access page content and structure for context
### Notion
The Notion integration enables AI SRE to:
* **Query databases** — Search and access Notion databases
* **Retrieve pages** — Access documentation and page content
* **Search knowledge** — Query your knowledge base for relevant information
* **Find content** — Search across your Notion workspace for context
## Next steps
# Observability
Source: https://docs-sre.lightrun.com/connectors/observability
Connect observability platforms to AI SRE
Connect AI SRE to your observability platforms to query metrics, logs, and traces during incident investigations.
The following connectors are available as beta features. You can request access to connectors based on your stack from the connectors screen in the web UI.
### Datadog
The Datadog integration enables AI SRE to:
* **Query metrics** — Access time-series data and performance metrics
* **Search logs** — Query and analyze log data across your services
* **Analyze traces** — Review distributed traces and APM data
* **Map services** — Understand service dependencies and relationships
### New Relic
The New Relic integration enables AI SRE to:
* **Access APM data** — Query application performance monitoring data
* **Monitor infrastructure** — Analyze infrastructure metrics and health
* **Search logs** — Query logs and events across your systems
* **Run custom queries** — Execute queries from your custom dashboards
### Prometheus
The Prometheus integration enables AI SRE to:
* **Query metrics** — Access time-series data using PromQL
* **Review alerts** — Analyze alerting rules and active alerts
* **Discover services** — Access service discovery information
* **Execute PromQL** — Run PromQL queries to gather metrics data
### Grafana
The Grafana integration enables AI SRE to:
* **Query dashboards** — Access Grafana dashboard data and panels
* **Query metrics** — Query time-series data from Grafana data sources
* **Search logs** — Access log data from Grafana Loki
* **Review alerts** — Access Grafana alerting rules and alert states
## Next steps
# Slack
Source: https://docs-sre.lightrun.com/connectors/slack-integration
Run investigations from Slack with @Lightrun AI SRE
Use **@Lightrun AI SRE** in Slack to investigate incidents without leaving your channels. Ask questions in natural language and get the same evidence-based analysis, root cause insights, and suggested actions you see in the web app—delivered directly in the channel.
## Overview
* **One workspace per organization** — Each organization connects one Slack workspace. Reconnecting replaces the existing connection.
* **Same capabilities as the web app** — Connected integrations, chat history, and model behavior are shared. Only the interface (Slack vs. browser) changes.
* **Scoped to mentions** — The app reads only messages where it is @-mentioned in a channel and replies in that channel.
## How to set up Slack
Log in to the AI SRE web app with your account.
In the sidebar, go to **Connectors** and select **Slack** to start the connection.
Choose the Slack workspace to connect and approve the **@Lightrun AI SRE** app when prompted.
After authorization, you’re redirected back to AI SRE. The workspace is connected and ready to use.
## Using Slack
In any channel where the app is installed, type **@Lightrun AI SRE** followed by your question or incident description. (Mentions work in channels only, not in threads.)
The app acknowledges your message and runs the same investigation pipeline as the web app. Responses stream into the channel as markdown.
Review evidence, root cause analysis, and suggested next steps in the channel. Reply and @-mention again in the channel for follow-up questions; the app uses the channel history as context.
Responses include the same evidence chains, confidence levels, and actionable recommendations as in the web UI. You can continue the conversation in the channel for deeper dives or clarifications.
## Troubleshooting
### Mentions get no response
The app must be in the channel to respond. Invite **@Lightrun AI SRE** to the channel if needed.
### Connection or install errors
Sign in to the AI SRE web app before starting the Slack connection. If the install flow fails, try again after signing in and ensure pop-ups or redirects aren’t blocked.
### Wrong workspace connected
Each organization can have only one connected Slack workspace. To use a different workspace, disconnect the current one from **Connectors** → **Slack** in the AI SRE web app, then connect the desired workspace.
## Disconnecting Slack
To remove the Slack connection:
1. In the AI SRE web app, go to **Connectors** → **Slack**.
2. Choose to disconnect the workspace.
3. Confirm. The app will no longer respond in Slack until you reconnect.
## Next steps
# FAQ
Source: https://docs-sre.lightrun.com/faq/index
Frequently asked questions
## What is AI SRE?
AI SRE is an AI-powered incident response assistant that helps support engineers, production engineers, on-call engineers, application engineers, and SREs investigate and resolve production incidents faster.
## How do I start an investigation?
Ask questions about incidents in natural language in the web UI or Slack. AI SRE gathers evidence from your tools, correlates data, and provides evidence-based insights with confidence levels.
## What do I need to get started?
* GitHub integration (required)
* Access to the AI SRE web application
* (Optional) Other connectors based on your stack
## How do I enable additional connectors?
Additional connectors are available as beta features. You can request access to connectors based on your stack from the connectors screen in the web UI.
## What types of incidents can AI SRE help with?
AI SRE helps investigate production incidents including service failures, performance degradation, errors, and outages. It correlates code changes, queries telemetry data, and builds evidence chains to identify root causes.
## What if AI SRE can't find the root cause?
AI SRE provides confidence levels with all findings. If confidence is low or no root cause is identified, it will indicate what data is missing or what additional investigation might help. You can refine your questions or connect additional data sources to improve results.
## How do I share investigation results?
Investigation results can be shared directly from the web UI or through Slack. Findings include evidence chains, confidence levels, and suggested actions that can be shared with your team.
## Can it make changes to my systems?
No. AI SRE is investigative only. It provides analysis and recommendations, but you make decisions and implement fixes.
## Is my data secure?
Yes. AI SRE uses enterprise-grade security with tenant isolation, encryption in transit and at rest, and read-only integrations. Your data is isolated per company and never used for AI model training.
## Security & Privacy
### Data isolation
Your data is isolated per company with dedicated storage and access controls. AI agent sandboxes are fully isolated per tenant.
### Access controls
* **Read-only integrations** — All integrations are configured without write permissions
* **Least-privilege access** — Permissions are scoped to the minimum required
* **Secure credential storage** — API keys are stored securely with tenant-specific isolation
### Data retention
Data retention is 14 days by default, but this is configurable.
### Source code protection
Source code is analyzed in real-time with no persistent storage. Only small code snippets from investigations are saved, and access is controlled via customer-managed OAuth tokens.
### AI model training
Your data is not used for AI model training. AI providers have zero data retention and do not train on customer inputs beyond processing API requests.
### Data protections
AI SRE uses industry-standard data protection practices including zero-retention contractual terms, Data Processing Agreements (DPAs), and prompt sanitization.
## How do I update my GitHub repository selection?
Go to **Connectors** → **GitHub** in the web UI to view connected repositories. To update, modify repository access in your GitHub App settings, and changes will sync automatically.
## Why isn't my connector working?
Check that:
* The connector is properly authenticated
* You have the required permissions
* The connector is enabled in the connectors screen
If issues persist, verify your API keys or OAuth tokens are valid and have the necessary permissions.
# GitHub Authentication
Source: https://docs-sre.lightrun.com/getting-started/github-auth
Connect GitHub repositories to AI SRE
Connect your GitHub repositories to enable AI SRE to access your codebase, track changes, and correlate deployments with incidents.
## Overview
When connected, AI SRE:
* **Accesses code** — Reviews code structure, dependencies, and recent changes
* **Tracks deployments** — Monitors commits, pull requests, and deployment timelines
* **Correlates changes** — Links code changes with incidents to identify root causes
* **Maps dependencies** — Understands service relationships and ownership
## Setup during onboarding
GitHub connection happens during onboarding:
1. **Sign in** — Log in to AI SRE
2. **Company setup** — If you're the first user, create your company account
3. **Connect GitHub** — The wizard prompts you to connect GitHub
4. **Authorize** — You'll be redirected to GitHub to authorize the AI SRE GitHub App
5. **Select organization** — Choose the GitHub organization to connect
6. **Select repositories** — Choose which repositories AI SRE can access
7. **Complete** — Return to AI SRE to finish setup
## Repository selection
Connect repositories that contain services you investigate during incidents. AI SRE uses these repositories to review code changes, track deployments, and correlate with incidents.
**Best practices:**
* Start with 1–2 critical production repositories
* Include repositories with active development and frequent deployments
* Add additional repositories as your needs grow
* Regularly review and update your connected repositories
## Troubleshooting
### Insufficient permissions
If you see an "Insufficient Permissions" error, you don't have the required permissions to install the GitHub App at the organization level.
**Solution:** Ask your GitHub organization administrator to install the app, or install it at the repository level if you have repository admin access.
### Repository not found
If a repository isn't accessible, verify:
* You have access to the repository
* The repository visibility settings allow your access level
* The repository exists and hasn't been deleted or renamed
### Authentication failed
If the OAuth flow fails:
* Try the authentication process again
* Clear your browser cache and cookies
* Ensure pop-up blockers aren't preventing the GitHub authorization page
## Managing connections
### View connected repositories
Navigate to **Connectors** → **GitHub** in the AI SRE interface to see which repositories are connected and their connection status.
### Update repository selection
To add or remove repositories:
1. Go to your GitHub App settings
2. Modify repository access permissions
3. Changes automatically sync to AI SRE
### Disconnect GitHub
To disconnect GitHub:
1. Go to **Connectors** → **GitHub** in AI SRE
2. Click **Disconnect**
3. Confirm the disconnection
**Note:** Disconnecting removes AI SRE's access to repository data. You'll need to reconnect to use GitHub-related features.
## Next steps
# Get Started
Source: https://docs-sre.lightrun.com/getting-started/overview
Set up AI SRE to start investigating production incidents
## Setup steps
Log in to the AI SRE web interface at `https://ai-sre.lightrun.com`.
During onboarding, authenticate with GitHub and select repositories to connect. See [GitHub Authentication](/getting-started/github-auth) for details.
Ask questions about incidents in natural language. AI SRE gathers evidence and provides insights.
**Note:** GitHub integration is required and is part of the onboarding flow.
## Next steps
# Quick Tour
Source: https://docs-sre.lightrun.com/getting-started/quick-tour
Overview of the AI SRE interface
The AI SRE interface uses a chat-based interaction model. Ask questions about incidents in natural language and receive evidence-based analysis with confidence levels.
## Using the chat interface
Type questions about incidents:
* "Why is checkout failing?"
* "What changed before this incident started?"
* "What's the impact of this error?"
AI SRE provides evidence-based analysis, confidence levels, and suggested diagnostic actions.
## Next steps
# On-Call Engineers
Source: https://docs-sre.lightrun.com/workflows/on-call-engineers
AI SRE workflow for On-Call Engineers
On-Call Engineers respond to alerts and incidents during their shifts. AI SRE helps understand noisy alerts, investigate with limited context, and respond effectively.
## Workflow stages
### Alert intake & triage
**Challenge:** Gets noisy alerts with limited context; scans logs and dashboards
**How AI SRE helps:**
* Provides context for noisy alerts
* Identifies important signals in alert noise
* Rapidly explains what alerts mean
* Helps prioritize alerts based on evidence
**Example:**
```
[Receives noisy alert]
You: "What's this alert about? Is it critical?"
AI SRE: [Analyzes alert, provides context, assesses severity]
```
### Scope & impact assessment
**Challenge:** Infers blast radius from service metrics; guesses severity
**How AI SRE helps:**
* Provides evidence-based blast radius
* Identifies affected services and dependencies
* Quantifies severity with evidence
* Replaces guesses with evidence
### Root cause investigation
**Challenge:** Greps logs; discovers missing fields; asks for redeploy with more logging
**How AI SRE helps:**
* Analyzes logs automatically
* Identifies what data is missing
* Gathers evidence from available sources
* Reduces manual log scanning
### Fix design
**Challenge:** Reviews fix for urgency, not correctness
**How AI SRE helps:**
* Reviews fixes for correctness, not just urgency
* Suggests fixes based on evidence
* Assesses fix risks
* Validates fix approach
### Deployment & verification
**Challenge:** Watches dashboards and error rates post-deploy
**How AI SRE helps:**
* Monitors system health automatically
* Verifies fixes are working
* Assesses impact reduction
* Identifies if fix didn't work
### Post-incident learning
**Challenge:** Moves on once alerts stop
**How AI SRE helps:**
* Documents investigation automatically
* Provides root cause summary
* Captures learnings from incident
* Retains investigation knowledge
## Key workflows
### Alert triage
1. Receive alert with limited context
2. Ask AI SRE to analyze alert
3. Get context and severity assessment
4. Prioritize based on evidence
5. Take appropriate action
### Quick investigation
1. Get alert or incident report
2. Ask AI SRE to investigate
3. Get evidence-based findings
4. Understand root cause
5. Take action
## Best practices
* Use AI SRE immediately when alerts come in
* Get context quickly before acting
* Verify fixes with AI SRE
* Document learnings before moving on
* Don't just move on once alerts stop
## Next steps
# Production Engineers
Source: https://docs-sre.lightrun.com/workflows/production-engineers
AI SRE workflow for Production Engineers
Production Engineers are often pulled into incidents late with little evidence. AI SRE helps quickly understand the situation, gather evidence, and contribute effectively to incident resolution.
## Workflow stages
### Alert intake & triage
**Challenge:** Often pulled in late with little evidence
**How AI SRE helps:**
* Provides immediate context about the incident
* Summarizes evidence gathered so far
* Explains current state of investigation
* Enables rapid onboarding to incident
### Scope & impact assessment
**Challenge:** Asked to validate impact without runtime data
**How AI SRE helps:**
* Provides impact assessment from available evidence
* Correlates available data to assess impact
* Identifies what runtime data is missing
* Supports impact validation with available data
### Root cause investigation
**Challenge:** Attempts to reproduce locally
**How AI SRE helps:**
* Analyzes production data directly
* Gathers evidence from production systems
* Provides guidance on reproduction if needed
* Focuses on production evidence, not just reproduction
### Fix design
**Challenge:** Designs fixes based on hypotheses, not real failure states
**How AI SRE helps:**
* Provides evidence of actual failure states
* Uses production data to understand failures
* Enables fix design based on evidence
* Validates fix design against real failure states
### Deployment & verification
**Challenge:** Lacks direct proof fix addressed the root cause
**How AI SRE helps:**
* Verifies fixes address root cause
* Provides evidence that fix worked
* Validates fix addresses root cause
* Provides confidence in fix effectiveness
### Post-incident learning
**Challenge:** Knowledge remains undocumented or tribal
**How AI SRE helps:**
* Documents investigation automatically
* Captures knowledge from investigation
* Creates shareable documentation
* Retains knowledge for future reference
## Key workflows
### Rapid onboarding
1. Get pulled into incident
2. Ask AI SRE for incident summary
3. Get evidence and current state
4. Understand investigation progress
5. Contribute effectively
### Evidence gathering
1. Identify what evidence is needed
2. Ask AI SRE to gather evidence
3. Get evidence from available sources
4. Fill evidence gaps
5. Contribute to investigation
## Best practices
* Use AI SRE immediately when pulled into incident
* Get context fast
* Gather evidence systematically
* Validate fixes with AI SRE
* Document knowledge for future
## Next steps
# SREs
Source: https://docs-sre.lightrun.com/workflows/sres
AI SRE workflow for Site Reliability Engineers
Site Reliability Engineers are responsible for system reliability, incident response, and ensuring services meet SLOs. AI SRE helps investigate incidents, perform root cause analysis, and improve system reliability.
## Workflow stages
### Alert intake & triage
**Challenge:** Reviews alert quality; suspects missing signals
**How AI SRE helps:**
* Identifies missing signals in alerts
* Evaluates alert quality and completeness
* Identifies what signals are missing
* Suggests improvements to alerting
### Scope & impact assessment
**Challenge:** Reconstructs dependencies from diagrams or tribal knowledge
**How AI SRE helps:**
* Analyzes dependencies from code and integrations
* Maps dependencies from evidence, not diagrams
* Identifies blast radius from system analysis
* Quantifies impact with evidence
### Root cause investigation
**Challenge:** Drives investigation but blocked by lack of evidence
**How AI SRE helps:**
* Gathers evidence from multiple sources
* Correlates data across systems
* Builds evidence chains systematically
* Identifies what evidence is missing
### Fix design
**Challenge:** Pushes for defensive fixes to reduce risk
**How AI SRE helps:**
* Suggests fixes based on evidence
* Recommends fixes that address root cause
* Assesses fix risks with evidence
* Validates fix approach
### Deployment & verification
**Challenge:** Monitors SLIs/SLOs, hoping metrics stabilize
**How AI SRE helps:**
* Monitors SLOs automatically
* Verifies fixes address root cause
* Validates fixes with evidence
* Provides confidence in resolution
### Post-incident learning
**Challenge:** Authors postmortems from partial data
**How AI SRE helps:**
* Provides complete investigation data
* Documents all evidence gathered
* Reconstructs complete timeline
* Enables comprehensive postmortems
## Key workflows
### Alert quality review
1. Review alerts with AI SRE
2. Identify missing signals
3. Assess alert quality
4. Improve alerting
5. Validate improvements
### Deep investigation
1. Start investigation with AI SRE
2. Gather evidence systematically
3. Build evidence chain
4. Identify root cause
5. Document findings
### SLO management
1. Monitor SLOs with AI SRE
2. Investigate SLO breaches
3. Identify root causes
4. Implement fixes
5. Validate SLO recovery
## Best practices
* Use AI SRE systematically for investigations
* Verify findings before acting
* Improve alerting based on AI SRE insights
* Use AI SRE for comprehensive documentation
* Build complete evidence chains
## Next steps
# Support Engineers
Source: https://docs-sre.lightrun.com/workflows/support-engineers
AI SRE workflow for Support Engineers
Support Engineers are often the first to interact with incidents from customer reports or monitoring systems. AI SRE helps correlate customer reports with system alerts and provides technical context for effective responses.
## Workflow stages
### Alert intake & triage
**Challenge:** Receives customer reports with vague symptoms; struggles to correlate tickets to alerts
**How AI SRE helps:**
* Correlates customer symptoms with system issues
* Maps customer tickets to relevant alerts
* Provides technical context from customer descriptions
* Assesses severity based on customer impact
**Example:**
```
Customer: "Checkout is slow"
You: "Why is checkout slow for customers?"
AI SRE: [Correlates with alerts, identifies performance issues]
```
### Scope & impact assessment
**Challenge:** Escalates based on customer complaints rather than system insight
**How AI SRE helps:**
* Provides evidence-based impact assessment
* Identifies affected services and dependencies
* Quantifies impact beyond customer complaints
* Provides technical evidence for escalation
**Example:**
```
You: "What's the impact of checkout being slow?"
AI SRE: [System-level impact: services, user %, business impact]
```
### Root cause investigation
**Challenge:** Relays findings manually between teams
**How AI SRE helps:**
* Performs investigation automatically
* Gathers evidence from multiple sources
* Provides shareable investigation results
* Findings can be shared directly with engineering teams
### Fix design
**Challenge:** Communicates assumptions back to customers
**How AI SRE helps:**
* Provides evidence-based explanations
* Translates technical findings clearly
* Indicates confidence levels
* Enables accurate status updates
### Deployment & verification
**Challenge:** Waits for confirmation to update customers
**How AI SRE helps:**
* Provides real-time system status
* Verifies when fixes are deployed and working
* Monitors impact reduction
* Enables timely customer communication
### Post-incident learning
**Challenge:** Writes support summaries with limited technical depth
**How AI SRE helps:**
* Provides technical details for summaries
* Documents investigation timeline
* Includes evidence chain
* Creates comprehensive summaries
## Key workflows
### Customer report correlation
1. Receive customer report with vague symptoms
2. Ask AI SRE to investigate based on customer description
3. Get technical correlation with system alerts
4. Map customer issue to specific system problems
5. Escalate with technical context
### Impact assessment
1. Get customer complaint
2. Ask AI SRE for system-level impact assessment
3. Understand blast radius and severity
4. Make informed escalation decision
5. Communicate impact to stakeholders
## Best practices
* Use AI SRE immediately when reports come in
* Correlate customer reports with alerts systematically
* Get technical context before escalating
* Provide technical evidence when escalating
* Use AI SRE findings for comprehensive documentation
## Next steps
# AI SRE Overview
Source: https://docs-sre.lightrun.com/working-with-ai-sre/overview
What AI SRE does and how it works
AI SRE is an AI-powered incident response assistant that helps support engineers, production engineers, on-call engineers, application engineers, and SREs investigate and resolve production incidents faster.
## Overview
AI SRE investigates incidents by:
* **Accessing code** — Reviews code changes, commits, and correlates them with incidents
* **Querying systems** — Gathers evidence from telemetry, infrastructure, and knowledge bases
* **Building evidence chains** — Correlates findings across systems to identify root causes
* **Delivering insights** — Provides evidence-based conclusions with confidence levels and suggested actions
## Benefits
* **Faster MTTR** — AI SRE does the investigation work, reducing investigation time from hours to minutes
* **Evidence-based** — Uses facts from your systems, not speculation—every finding is backed by data from your code, logs, and metrics
* **Works with your tools** — No data migration required; connects to your existing stack
* **Confidence levels** — Clearly states certainty and what's missing, so you know when to verify
## Architecture
```mermaid theme={null}
graph LR
A[You ask question] --> B[AI SRE gathers evidence]
B --> C[Correlates data]
C --> D[Provides findings]
D --> E[You take action]
```
1. **You ask** — Type questions about incidents in natural language
2. **AI SRE works** — Queries code repositories, scans logs and metrics, reviews recent changes, and correlates data across your stack
3. **AI SRE correlates** — Links findings across systems, builds evidence chains, and identifies root causes
4. **AI SRE delivers** — Provides evidence-based conclusions with confidence levels and actionable next steps
5. **You resolve** — Use insights to fix incidents faster
## Use cases
* **"Explain the latest code change and its impact"** — AI SRE reviews recent deployments and code changes, correlating with current system state to explain what changed and how it affects your services
* **"Analyze this alert and summarize what's happening"** — AI SRE queries logs, metrics, and traces to understand the alert context, identifies affected services, and provides a clear summary of what's broken
* **"What caused this incident?"** — AI SRE investigates by reviewing code changes, correlating with incident timeline, querying telemetry data, and building evidence chains to identify the root cause
* **"Which team owns this service?"** — AI SRE searches code repositories, reviews ownership patterns, and identifies team information from your knowledge base
## Data Security & Privacy
Your data is isolated per company, encrypted in transit and at rest, and never used for AI model training. Integrations use read-only access. See the [Security & Privacy](/faq#security--privacy) section for details.
## Limitations
* **Read-only** — AI SRE cannot make changes to your systems
* **Data availability** — Insights depend on connected integrations and data quality
* **Confidence levels** — Findings include confidence levels; always verify critical decisions
* **Tool expertise** — Works best with well-configured integrations and monitoring
## Get started
Set up AI SRE
Learn workflows