# Sentrely — Full Content Dump for LLM Crawlers > Managed control plane for AI agents. Policy enforcement, audit trails, and human-in-the-loop approvals for the Claude, Cursor, Codex, and OpenAI agents your team builds — fully managed, no infra to set up. Last generated: 2026-07-26T18:05:55.643Z Canonical site: https://sentrely.com --- ## Product summary Sentrely is a B2B SaaS control plane for autonomous AI agents. It sits as a gateway between the agents your team builds (Claude Code, Cursor agents, OpenAI Codex, etc.) and the resources they touch (AWS, GitHub, Stripe, internal APIs). Sentrely enforces YAML-defined RBAC policies on every request, records an immutable audit trail of every tool call, and routes risky operations to humans for approval via Slack or Telegram. Pricing: Free 14-day trial · Team $199/mo · Business $999/mo · Enterprise custom (VPC deployment). Connection methods: API key (Anthropic, OpenAI, Mistral) and OAuth (Cursor Pro/Business, ChatGPT Plus, Claude Pro). Bring your own subscription — no credits, no seat fees beyond included counts. --- ## Comparison pages ### Sentrely vs AgentOps **URL:** https://sentrely.com/compare/vs-agentops **Verdict:** AgentOps monitors your agents. Sentrely controls them. **Feature comparison:** | Feature | Sentrely | AgentOps | |---|---|---| | Trace agent sessions | ✓ | ✓ | | Real-time policy enforcement | ✓ | ✕ | | Block actions before they happen | ✓ | ✕ | | Human-in-the-loop approvals | ✓ | ✕ | | Per-agent RBAC policies | ✓ | ✕ | | Multi-provider failover | ✓ | ✕ | | LLM cost analytics dashboard | Limited | ✓ | | Replay session debugging | ✓ | ✓ | **FAQ:** **Q: Is Sentrely an AgentOps alternative?** Partially. AgentOps is best for monitoring and post-hoc analytics — what your agents did, how much they cost, where they failed. Sentrely is a control plane — it monitors and controls what your agents do in real time. If you only need observability, AgentOps is great. If you need to gate risky actions on human approval, Sentrely is purpose-built for that. **Q: Can I use both together?** Yes. Sentrely produces structured audit logs that you can stream into AgentOps for richer analytics. Many teams use Sentrely for runtime governance and AgentOps for cost analysis. **Q: Why pick Sentrely over AgentOps?** If your agents are touching production systems (AWS, git, payment systems, customer data), you need more than observability — you need policy enforcement and approval gating. AgentOps will tell you that your agent leaked data; Sentrely prevents it from happening in the first place. **Q: Does Sentrely have cost analytics like AgentOps?** Sentrely has per-agent and per-project token budgets, spend alerts, and rate limits. For deep cost analysis with attribution and trends, AgentOps has more sophisticated dashboards. The two complement each other. **Q: Does Sentrely work with the same agent frameworks as AgentOps?** Sentrely is framework-agnostic — works with any agent that makes HTTP calls. AgentOps has SDK-level integration with specific frameworks (CrewAI, AutoGen, LangChain). Different integration models, similar coverage. **Article:** AgentOps is an observability platform for AI agents — it helps you understand what your agents are doing, debug problems, and analyze behavior over time. Sentrely is a control plane — it enforces what agents are allowed to do and gives you mechanisms to intervene in real time. Both create visibility into agent behavior. Only one acts on that visibility. ## AgentOps' Strengths AgentOps focuses on monitoring and debugging: - **Session replay.** See a complete recording of an agent session — every step, every tool call, every decision. - **Error tracking.** Catch and catalog agent errors with context. - **Performance analytics.** Latency, token usage, cost breakdown per session. - **LLM call tracing.** Full visibility into what was sent to and received from the model. - **Multi-framework support.** Works with LangChain, CrewAI, AutoGen, and custom implementations. If you need to understand why an agent is behaving unexpectedly, AgentOps gives you the diagnostic tools. ## What AgentOps Doesn't Do - **Policy enforcement.** AgentOps doesn't intercept requests to enforce access policies. It records what happened. - **RBAC.** No per-agent permission scoping. - **Real-time intervention.** By the time you see an issue in AgentOps, it's already happened. - **Approval gates.** No mechanism to pause an agent pending human review. - **Cost controls.** No token budgets or session termination based on spending. - **Compliance audit trail.** AgentOps traces are excellent for debugging; the format and retention aren't designed for SOC 2 or HIPAA compliance evidence. ## The Monitoring vs. Control Distinction This is the clearest way to understand the difference: **Monitoring asks:** What did my agent do? **Control asks:** What is my agent allowed to do, and what happens when it tries to do something it shouldn't? AgentOps is excellent at the first question. Sentrely answers both. | Capability | AgentOps | Sentrely | |---|---|---| | Session recording | Full replay | Action-level log | | LLM call tracing | Detailed | Summary | | Error tracking | Excellent | Alert-based | | Performance analytics | Strong | Basic | | Policy enforcement | None | Core feature | | RBAC | None | Per-agent policies | | Approval gates | None | Slack/Telegram/dashboard | | Cost controls | Visibility only | Budget enforcement | | Real-time kill switch | No | Yes | | Compliance audit trail | Partial | Yes | ## Using Both These tools complement each other well. Teams often use AgentOps during development — debugging agent behavior, improving prompts, finding edge cases — and Sentrely in production — enforcing policies, managing access, generating compliance evidence. If you're choosing one for a production deployment where control and compliance are primary concerns, Sentrely addresses the operational risk layer. If you're choosing one for development where debugging and quality are primary concerns, AgentOps is stronger. For teams running Claude Code agents specifically, the distinction is stark: you need the control plane first (before anything goes to production), and the observability layer is a valuable addition. --- ### Sentrely vs Anthropic Managed Agents **URL:** https://sentrely.com/compare/vs-anthropic-managed-agents **Verdict:** Anthropic Managed Agents handles execution infrastructure so you don't have to. Sentrely adds the governance layer — per-agent RBAC, human approvals, compliance audit trails — that the official service doesn't provide. **Feature comparison:** | Feature | Sentrely | Anthropic Managed Agents | |---|---|---| | Provider-agnostic (works with non-Anthropic models) | ✓ | ✕ | | Per-agent RBAC policies | ✓ | Limited | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | ✕ | | Independent audit trail (not tied to vendor) | ✓ | ✕ | | Multi-provider failover | ✓ | ✕ | | Run on your own infra / VPC | ✓ | ✕ | | Native Claude execution infrastructure | ✕ | ✓ | | Anthropic-managed scaling | ✕ | ✓ | **FAQ:** **Q: What's the difference between Sentrely and Anthropic Managed Agents?** Anthropic Managed Agents is Anthropic's official execution service — they run your Claude agents on their infrastructure. Sentrely is a vendor-neutral control plane that sits in front of any agent (Claude, Cursor, Codex, GPT) and enforces your policies, audits every action, and gates risky ops on human approval. Anthropic handles the runtime; Sentrely handles the governance. **Q: Can I use Sentrely with Anthropic Managed Agents?** Yes. Sentrely sits as a gateway between your agents and the outside world. You can run agents on Anthropic's managed runtime and route every API call, tool invocation, and resource access through Sentrely's policy engine. **Q: Why pick Sentrely over Anthropic's official service?** Three reasons: (1) Provider neutrality — Sentrely works with Claude, OpenAI, Cursor, and Codex equally, so you're not locked to one vendor. (2) Independent audit trail — your compliance data isn't sitting inside Anthropic's logs. (3) Custom RBAC — define exactly which AWS actions, git repos, and tools each agent can touch, with approvals routed to your Slack. **Q: Is Sentrely as scalable as Anthropic's managed runtime?** Sentrely is a gateway, not a runtime. We scale to handle policy decisions on every request (microseconds of overhead). The actual agent execution runs wherever you want — Anthropic's runtime, your AWS account, a VPC, or a developer laptop. **Q: What if I'm already using Anthropic's managed agent service?** Add Sentrely as a layer in front. Your agents continue to run on Anthropic's infra; Sentrely captures every tool call and adds policy enforcement. No code changes needed. **Article:** ## What Anthropic Managed Agents Is Anthropic's Managed Agents is an official hosted execution service for Claude agents. At $0.08 per session hour, it handles the infrastructure complexity of running autonomous agents: sandboxed code execution, session state management, credential handling, tool definition, checkpointing, and end-to-end tracing. The pitch is speed: go from idea to production agent in days instead of months. It's the official Anthropic answer to "I just want to run Claude agents without building infrastructure." ## What Sentrely Is Sentrely is an independent control plane for Claude Code agents. It provides the governance layer: per-agent RBAC policies that restrict what agents can access, human-in-the-loop approval workflows via Slack and Telegram, immutable audit trails designed for compliance auditors, and cost controls that prevent runaway agent spend. Sentrely isn't competing with Anthropic's execution layer — it's the governance layer that sits on top of it. ## The Gap in the Official Service Anthropic Managed Agents excels at execution infrastructure. What it doesn't provide: **Per-agent access control.** The managed service runs your agents with the permissions you configure at setup. There's no policy engine that says "agent A can read S3 bucket X but not Y" or "agent B can push to feature branches but requires approval for main." All agents run under the same credential scope. **Human approval gates.** If your agent wants to delete a production database table, the managed service will execute that operation if the agent has the permission. There's no mechanism to pause and route that specific decision to a human in Slack before it runs. **Compliance audit trail.** Anthropic's tracing gives you end-to-end session visibility for debugging. That's different from the structured, immutable, field-level audit logs that SOC 2 and HIPAA auditors require. A compliance-grade audit log isn't an afterthought — it's a specific artifact with specific fields and retention requirements. **Multi-agent isolation.** When you're running 10 agents on the same project, the managed service doesn't enforce that agent A can't access agent B's resources. Sentrely enforces per-agent, per-resource isolation by policy. ## The Comparison | Dimension | Anthropic Managed Agents | Sentrely | |---|---|---| | Session infrastructure | Yes — handled entirely | You run your own agents | | Sandboxed code execution | Yes | No | | Session state / checkpointing | Yes | No | | Credential management | Yes — built-in | Via Gateway credential vending | | Per-agent RBAC | No | Yes — per-resource policies | | Human approval gates | No | Yes — Slack + Telegram | | Immutable audit trail | Debugging traces | Compliance-grade logs | | SOC 2 / HIPAA evidence | No | Yes | | Multi-agent isolation | No | Yes | | Token budgets (per session) | Session pricing | Hard per-session budget caps | | Kill switch | Session termination | Yes + fleet management | | Vendor lock-in | Anthropic only | Claude-focused, portable design | | Pricing | $0.08/session hour + API costs | Starter $49/mo flat | | Self-hosted / VPC option | No | Enterprise tier | ## When to Use Anthropic Managed Agents The official service is the right choice when: - You want the absolute fastest path to a running Claude agent - Session infrastructure (sandboxing, state, checkpointing) is your main concern - Governance requirements are light — no compliance audits, no human approval workflows - You want first-party support from Anthropic on execution issues ## When to Use Sentrely Sentrely is the right choice when: - Individual agents need different permission scopes - Operations like production deploys require human sign-off - You need SOC 2 or HIPAA audit evidence for AI agent operations - Per-agent cost attribution and hard caps matter - You don't want a single vendor (Anthropic) controlling both your model and your governance infrastructure ## Can You Use Both? Yes, and this is actually a natural architecture. Anthropic Managed Agents handles the execution layer (sandboxing, session state, credential vending for Anthropic services). Sentrely sits above it handling governance (what agents can access, what requires approval, what gets logged for compliance). If Anthropic's execution service meets your infrastructure needs and Sentrely meets your governance needs, using both is a reasonable production architecture. The risk to monitor: as Anthropic adds governance features to their managed service, the overlap will grow. ## The Vendor Independence Question One consideration that doesn't appear in feature tables: when your model provider and your governance layer are the same company, policy changes in that relationship affect both. Sentrely is an independent control plane — changes to Anthropic's pricing, terms, or product direction don't simultaneously affect your governance infrastructure. For organizations where auditability and independence matter, keeping execution and governance in separate vendor relationships is a meaningful architectural choice. --- ### Sentrely vs Microsoft Copilot Studio **URL:** https://sentrely.com/compare/vs-copilot **Verdict:** Copilot Studio builds AI assistants for Microsoft-ecosystem users. Sentrely governs autonomous Claude Code agents running against your engineering infrastructure. Different products for different problems. **Feature comparison:** | Feature | Sentrely | Microsoft Copilot Studio | |---|---|---| | Provider-agnostic (Claude, OpenAI, Cursor) | ✓ | ✕ | | Code-first agent governance | ✓ | Limited | | Per-agent RBAC for AWS, git, tools | ✓ | ✕ | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | Teams only | | Audit trail streamed to S3 / your bucket | ✓ | Microsoft Purview | | No-code agent builder | ✕ | ✓ | | Microsoft 365 / Teams native integration | ✕ | ✓ | | Power Platform connectors (1000+) | ✕ | ✓ | **FAQ:** **Q: Is Sentrely a Copilot Studio alternative?** Only partially. Copilot Studio is for business users building AI assistants inside Microsoft 365 — workflows that live in Teams, SharePoint, Power Platform. Sentrely is for engineering teams running autonomous agents (Claude Code, Cursor, Codex) against production infrastructure. If your AI work happens in Word and Excel, Copilot Studio. If it happens in your IDE and CI pipeline, Sentrely. **Q: Can I use both?** Yes. Many enterprises use Copilot Studio for business-side automation and Sentrely to govern the Claude/Cursor agents their engineers run. They sit in different parts of the company stack and don't conflict. **Q: Does Sentrely integrate with Microsoft Teams?** Yes — approval requests can be routed to Teams (in addition to Slack and Telegram), and audit logs can be exported for Microsoft Purview ingestion if your compliance team requires it. **Q: Why isn't Copilot Studio enough for my engineering agents?** Copilot Studio is designed for low-code workflows and Microsoft-ecosystem data. It doesn't have the policy granularity to control what a Claude agent does to your AWS account, your GitHub org, or your Stripe API. Sentrely is purpose-built for that level of control. **Q: What if I'm an enterprise Microsoft customer?** Sentrely's Enterprise plan supports SSO via Azure AD / Entra ID, integrates with Microsoft Purview for compliance, and can deploy in your Azure VPC. Most large enterprises run both — Copilot Studio for business processes and Sentrely for engineering agents. **Article:** ## What Microsoft Copilot Studio Is Microsoft Copilot Studio is a low-code platform for building custom AI copilots that extend Microsoft 365. You connect it to your data sources (SharePoint, Dynamics, external APIs), define conversational flows, and deploy across Teams, web chat, mobile, or your own applications. It's built for line-of-business users and IT teams who want to create AI assistants without writing code — scheduling meetings, answering HR questions, looking up customer records, triggering workflows. Copilot Studio is deeply embedded in the Microsoft ecosystem. If your organization runs on M365 and Azure, it's the natural path to AI automation for non-technical users. ## What Sentrely Is Sentrely is a managed control plane for Claude Code agents — autonomous AI agents that read and write code, access AWS infrastructure, push to git repositories, call external APIs, and make complex engineering decisions. These aren't chatbots responding to employee questions. They're agents running overnight, unattended, with real credentials and real blast radius. Sentrely governs what those agents can do: per-agent RBAC policies, human approval gates for risky operations, immutable audit trails for compliance, and cost controls to prevent runaway API spend. ## Different Buyers, Different Problems Copilot Studio and Sentrely serve fundamentally different buyers: **Copilot Studio buyer:** IT leader, line-of-business manager, or citizen developer who wants to automate workflows and create AI assistants for Microsoft-platform users. Building an HR chatbot, a customer service bot, or an internal knowledge assistant. **Sentrely buyer:** Engineering team running Claude Code agents for automated code review, CI/CD pipelines, data processing, or infrastructure management — and needing the security controls, audit trails, and governance that production systems require. | Dimension | Copilot Studio | Sentrely | |---|---|---| | Primary use case | Conversational AI assistants | Autonomous agent governance | | Target user | IT admin, business analyst | Engineering team, DevOps | | Ecosystem | Microsoft 365, Azure, Power Platform | Claude Code, AWS, git, any API | | Agent autonomy | Low — structured conversation flows | High — fully autonomous operations | | Code/git integration | Limited | Core feature | | Per-agent RBAC | Workspace-level permissions | Per-agent, per-resource policies | | Human approval gates | Manual workflow triggers | Automatic Slack/Telegram gates | | Audit trail | Microsoft Purview integration | Purpose-built compliance evidence | | Infrastructure access | Microsoft services primarily | Any AWS/cloud/git/API | | Non-Microsoft LLM support | Limited | Claude-native | | Self-hosted / VPC option | Azure deployment | VPC Enterprise tier | | Pricing | Power Platform licensing (~$200+/user/mo) | Starter $49/mo | ## The Autonomy Gap The key distinction is autonomy. Copilot Studio creates AI assistants that respond to user requests — someone asks a question, the copilot answers. Sentrely governs agents that act without being asked — running overnight, executing multi-hour tasks, making dozens of tool calls per minute. Autonomous action requires different governance: - A chatbot that answers HR questions doesn't need a kill switch. An agent deploying to production does. - A conversational AI doesn't need per-resource access policies. An agent touching customer databases does. - A bot answering employee questions doesn't produce a compliance audit trail. An agent processing regulated data must. If you're building AI assistants for internal users on Microsoft 365, Copilot Studio is purpose-built for that. If you're deploying Claude Code agents to do autonomous engineering work, you need a control plane designed for that level of autonomy and risk. ## A Word on GitHub Copilot GitHub Copilot (the code completion and chat tool in VSCode/JetBrains) is different from Copilot Studio and not a competitor to Sentrely. GitHub Copilot is a developer productivity tool. Sentrely governs autonomous Claude Code agents. The two can coexist — developers use GitHub Copilot day-to-day, while automated Claude agents run governed pipelines through Sentrely. --- ### Sentrely vs CrewAI **URL:** https://sentrely.com/compare/vs-crewai **Verdict:** CrewAI builds and runs your agents. Sentrely governs the agents you build. They solve different problems — and many teams need both. **Feature comparison:** | Feature | Sentrely | CrewAI | |---|---|---| | Multi-agent framework / orchestration | Gateway-level | ✓ | | Policy-based RBAC for agent actions | ✓ | ✕ | | Audit trail of every tool call | ✓ | ✕ | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | ✕ | | Multi-provider model routing & failover | ✓ | Limited | | Cost / token budgets per agent | ✓ | ✕ | | Build and define agent crews | ✕ | ✓ | | Pre-built agent role templates | ✕ | ✓ | **FAQ:** **Q: Is Sentrely a CrewAI alternative?** No — CrewAI is a Python framework for building multi-agent systems. Sentrely is the gateway you put in front of any agent (built with CrewAI, LangChain, raw Python, or no framework at all) to enforce policies, audit actions, and gate risky operations on human approval. Most teams running CrewAI in production add Sentrely as the governance layer. **Q: Why use Sentrely with CrewAI?** CrewAI is great at orchestrating agent collaboration but has no policy enforcement, no audit trail, and no approval gating. The moment your CrewAI agents touch production infrastructure (AWS, git, Stripe), you need Sentrely to control what they're allowed to do and prove what they did. **Q: Does Sentrely require me to use CrewAI?** No — Sentrely is framework-agnostic. It sits at the network layer. Whether your agents are written with CrewAI, LangChain, AutoGen, or just raw Anthropic SDK calls, Sentrely intercepts every API call and tool invocation. **Q: How does Sentrely's multi-agent support compare to CrewAI?** Different layers. CrewAI orchestrates collaboration between agents (who does what, in what order). Sentrely tracks live sessions across multiple agents and routes A2A messages, but it doesn't define your agent crew structure. They're complementary. **Q: What about open-source CrewAI vs Sentrely?** CrewAI is open-source and free. Sentrely's gateway core is also open-source — managed Sentrely adds the dashboard, audit retention, SSO, multi-tenancy. Both products fit naturally with open-source workflows. **Article:** ## What CrewAI Is CrewAI is a multi-agent orchestration framework and platform. You define crews of agents — each with a role, tools, and goal — and CrewAI handles the execution: task delegation, agent communication, workflow sequencing, and results aggregation. It's one of the most popular agent frameworks in the world, used across the Fortune 500, and available as open-source (CrewAI OSS) or a managed cloud platform (CrewAI AMP). CrewAI excels at defining what agents do: their personas, their tools, their collaboration patterns. It's excellent at building the agent system. ## What Sentrely Is Sentrely is a managed control plane for Claude Code agents. It governs what agents are allowed to do in production: which AWS resources they can access, which git branches they can push to, whether destructive operations require human approval, how much they can spend per session. It's the security and compliance layer that sits above the execution framework. Sentrely doesn't build agents. It governs them. ## The Core Difference CrewAI and Sentrely operate at different layers of the stack: **CrewAI (execution layer):** Defines agents, assigns tasks, coordinates multi-agent workflows, returns results. **Sentrely (governance layer):** Enforces what agents can access, logs every action, gates risky operations on human approval, manages costs, provides the audit trail. A useful analogy: Kubernetes schedules and runs your containers. AWS IAM controls what those containers can access. CrewAI is closer to the scheduler; Sentrely is closer to the policy engine. ## What Each Does That the Other Doesn't | Capability | CrewAI | Sentrely | |---|---|---| | Define agent roles and personas | Yes — core feature | No | | Orchestrate multi-step agent workflows | Yes — core feature | No | | Task delegation between agents | Yes | Via A2A messaging | | Visual agent builder | Yes (CrewAI AMP) | No | | Per-agent RBAC policies | No | Yes — per-agent, per-resource | | Human approval gates (Slack/Telegram) | No | Yes | | Immutable audit trail | No | Yes | | Token budgets per session | No | Yes | | Session kill switch | No | Yes | | SOC 2 / HIPAA compliance evidence | No | Yes | | AWS / git / cloud resource governance | No | Yes | ## When to Use Each **Use CrewAI when** you need to define complex multi-agent workflows with specialized roles, build agent pipelines without writing orchestration code from scratch, or give non-technical stakeholders a visual interface to build and run agents. **Use Sentrely when** your agents touch real production systems (git, AWS, customer data, APIs), you need to pass a compliance audit, you need human sign-off on destructive or high-risk operations, or you need to know exactly what each agent did and when. **Use both when** you're running a CrewAI-built workflow in production. CrewAI handles the orchestration; Sentrely governs what each agent is allowed to touch. This is the most common production setup — frameworks handle complexity, control planes handle safety. ## The Production Problem CrewAI is excellent for building and demonstrating agents. The gap it doesn't fill: what happens when your agents run overnight, unattended, against real infrastructure? - Which S3 buckets can the research agent read? Which can it write? - If the deployment agent wants to push to main at 3am, who approves that? - If an agent loop runs for 2 hours and burns $400, what stops it? - When security asks for every action taken by `agent-3` last Tuesday, what do you show them? These questions aren't answered by an orchestration framework. They're answered by a control plane. Sentrely is built for exactly this layer. --- ### Sentrely vs Cursor **URL:** https://sentrely.com/compare/vs-cursor **Verdict:** Cursor is an AI-first IDE for individual developers. Sentrely governs the Cursor agents your team is already running. **Feature comparison:** | Feature | Sentrely | Cursor | |---|---|---| | AI-powered code editor (chat, tab-complete, agent mode) | ✕ | ✓ | | Routes Cursor agents through a control plane | ✓ | ✕ | | Policy-based RBAC for what agents can touch | ✓ | ✕ | | Full audit trail of every tool call & API request | ✓ | ✕ | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | ✕ | | Multi-provider model routing (Claude, GPT, etc.) | ✓ | ✓ | | Bring-your-own-subscription via OAuth | ✓ | ✓ | | Cost / token budgets per project | ✓ | Limited | | Multi-agent orchestration & A2A messaging | ✓ | ✕ | | Meant to replace your IDE | ✕ | ✓ | **FAQ:** **Q: Is Sentrely a Cursor alternative?** No — Sentrely doesn't replace Cursor. Cursor is the IDE your developers open every day. Sentrely is the control plane sitting behind Cursor, governing what Cursor's agent mode can do once it leaves the editor (push to git, run shell commands, hit AWS, etc.). They're complementary. **Q: Why do teams use Sentrely with Cursor?** Cursor agent mode is powerful but unaudited — it'll happily run any shell command or git push it thinks is right. For teams shipping to production, that's risky. Sentrely lets you define policies (e.g. 'never push to main without approval'), keeps an immutable audit log of every agent action, and gates risky operations on human approval delivered to Slack. **Q: Does Sentrely work with Cursor's existing OAuth?** Yes. Cursor agents authenticate to Sentrely via OAuth. Each developer signs in with their existing Cursor Pro account — no extra API keys needed. Sentrely then enforces team-level policies on every request from any logged-in Cursor user. **Q: How is Sentrely different from Cursor's Business plan?** Cursor Business adds team-level admin features inside the editor (SSO, usage analytics, privacy controls). Sentrely controls what agents do outside the editor — actual policy enforcement on AWS calls, git operations, and external APIs. Most teams use both: Cursor Business for editor management, Sentrely for runtime governance. **Q: Can I see what every Cursor developer's agents are doing?** Yes. Sentrely's dashboard shows live agent sessions, every tool call, and full conversation history per agent. You can filter by user, project, or action type, and replay any session. The audit log streams to your own S3 bucket for compliance. **Q: Do I need Sentrely if I'm a solo Cursor user?** Probably not. Sentrely's value comes from team-level governance — policies, approvals, audit trails. If you're solo and just want a great AI editor, Cursor alone is the right answer. Once you're on a team of 3+ developers running Cursor agent mode against shared infrastructure, Sentrely becomes worth it. **Q: What does Sentrely cost on top of Cursor?** Sentrely is per-team, not per-developer: $199/mo for the Team plan (10 agents, 2 admin seats), $999/mo Business (20 agents, 10 seats, SSO/SCIM). Cursor Pro is $20/user/month. So a 10-developer team running both pays roughly $200 (Cursor) + $199 (Sentrely Team) = $399/mo. **Article:** Cursor and Sentrely solve completely different problems. Cursor is an AI-first IDE — a fork of VS Code that gives developers Tab-completion, in-line chat, and agent mode for refactoring code. Sentrely is the control plane that sits *behind* tools like Cursor, governing what its agents are allowed to do once they leave the editor. If you're a solo developer writing code, you want Cursor. If you're a team running Cursor (or any other AI agent) against production infrastructure, you want both. ## What Cursor Is Cursor is the editor your developers actually open. It provides: - **Tab completion** that's aware of your whole repo - **Inline chat** ("⌘K") to edit blocks of code - **Agent mode** — autonomous multi-file refactors and bug fixes - **Composer** — repo-aware code generation - **MCP support** — connect Cursor to external tools Cursor lives on a developer's laptop. It writes code. It can run shell commands. It can hit your APIs. It's incredible for individual productivity. ## What Sentrely Is Sentrely is the gateway between your AI agents (including Cursor's agent mode) and the outside world. It provides: - **Policy-based RBAC** — which AWS actions, git repos, and tools each agent can touch - **Audit trail** — every tool call, approval, and API request, logged automatically - **Human-in-the-loop approvals** — risky operations gate on your approval, delivered via Slack or Telegram - **Multi-provider failover** — define a backup model so an outage at Claude doesn't stop your agents - **Cost controls** — per-project token budgets, rate limits, spend alerts Sentrely doesn't write code. It makes sure the agents that *do* write code don't push to `main` without approval, leak secrets, or burn $4.20 in tokens looping on a 429. ## When You Need Both This is the most common case for engineering teams that have adopted Cursor: 1. Developers use Cursor day-to-day for editing code. 2. Cursor's agent mode connects to Sentrely as the gateway URL. 3. Every shell command, AWS call, and tool invocation Cursor agents make routes through Sentrely's policy engine. 4. Risky operations (push to main, delete production data, rotate secrets) gate on approval in Slack. 5. Every agent run is recorded in an immutable audit log streamed to your S3. The result: your developers keep the productivity of Cursor agent mode, your security team gets the controls they need, and your audit team gets compliance evidence by default. ## Honest Trade-offs **Pick Cursor alone** if you're a solo dev or small startup with no production infrastructure to protect. The IDE experience is excellent on its own. **Pick Sentrely alone** if your agents aren't running through Cursor — you have custom Claude Code or OpenAI Codex agents in your CI/CD pipeline, on a server, or in a Lambda. **Pick both** if you're a team of 5+ developers using Cursor agent mode against shared infrastructure. Cursor accelerates writing; Sentrely keeps the writing safe. ## Pricing - **Cursor Pro**: $20/month per user. Cursor Business: $40/user/month. - **Sentrely**: $199/month for a team (10 agents, 2 seats). $999/month for business (20 agents, 10 seats, SSO/SCIM). These aren't competing line items — Cursor is per-developer, Sentrely is per-team. Most teams running both end up with Cursor Pro on every developer plus one Sentrely Team or Business plan. ## The Bottom Line Cursor makes your developers faster. Sentrely makes the agents your developers spawn safe to run in production. Use them together — Cursor in front of the keyboard, Sentrely behind every API call. --- ### Sentrely vs DIY / Build Your Own Gateway **URL:** https://sentrely.com/compare/vs-diy-gateway **Verdict:** You can build it. The question is whether you should — and what it will really cost. **Feature comparison:** | Feature | Sentrely | DIY / Build Your Own Gateway | |---|---|---| | Time to first deploy | 10 minutes | 3-6 months | | Engineering FTE cost | $0 | $200k+/yr | | Policy DSL maintained for you | ✓ | ✕ | | Production audit log infrastructure | ✓ | Build & maintain | | Multi-provider failover logic | ✓ | Build yourself | | SSO / SCIM | ✓ | Build yourself | | Compliance attestations (SOC2) | ✓ | $50-150k audit | | Total control over implementation | ✕ | ✓ | | Open-source self-host option | ✓ | ✓ | **FAQ:** **Q: How long does it take to build an AI agent control plane?** A minimum-viable version (basic auth, request logging, token rate limiting) is 4-8 weeks of senior engineering work. A production-grade version with policy DSL, approval workflows, multi-provider failover, audit retention, and compliance evidence is 3-6 months. Maintaining it indefinitely is roughly 0.25-0.5 of a senior engineer. **Q: What does DIY actually cost?** Conservative estimate for production-grade: ~$200-300k in initial engineering ($150k senior eng × 4-6 months at fully-loaded cost) + $50-100k/year in ongoing maintenance + $50-150k for SOC2 audit if you need compliance. Sentrely's Business plan is $11,988/year and includes all three. **Q: Should I ever build my own?** Yes, in two cases: (1) Your security/compliance posture forbids any third-party in the request path — air-gapped or sovereign clouds. (2) You have unique requirements no off-the-shelf gateway covers (e.g. custom policy languages tied to internal IAM). Both are rare for typical engineering teams. **Q: Can I start with Sentrely and migrate to DIY later?** Yes. Sentrely's gateway is open-source — the same code we run is on GitHub. If you ever decide to self-host or fork it, you can lift the gateway into your VPC and bring your existing policy YAML with you. **Q: What about the open-source version of Sentrely?** The gateway core is open-source and self-hostable. The managed product adds the dashboard, audit retention, multi-tenancy, SSO, and 24/7 ops. If you want to start free, self-host the OSS gateway. If you want to scale, move to managed. **Article:** Every engineer's first reaction to "you need a control plane for your Claude agents" is "I can build that." And you can. The question is whether you should — and what that decision actually costs. ## What "Build Your Own" Actually Requires A minimal viable gateway for Claude agents needs these components. Here's an honest estimate of the engineering time each takes: **Proxy layer** — intercept and route all agent API calls: 2-3 days **Authentication and agent identity** — unique credentials per agent, token management: 3-5 days **Policy engine** — load YAML policies, evaluate rules against requests, enforce deny-by-default: 1-2 weeks **Audit logging** — structured logging, immutable storage, queryable retrieval: 1 week **Approval workflows** — trigger on policy gates, route to Slack/Telegram, handle responses, timeout and retry: 1-2 weeks **Cost tracking** — token budgets per agent/session, alert thresholds, automatic session termination: 1 week **Session management** — start/stop/kill, session state, heartbeats: 3-5 days **Dashboard** — live session view, audit log search, policy editor: 2-3 weeks **Total first version: 2-3 months** of a senior engineer's time. That's before testing, documentation, and the edge cases you'll discover in production. ## The Hidden Costs The build estimate above is for the first version. What gets underestimated: **Ongoing maintenance.** The policy engine needs updates as your access patterns change. The audit log format evolves as you add integrations. The Slack integration breaks when Slack changes their API. Someone owns this. **Security issues.** A gateway that's handling credentials and access controls is a security-sensitive system. When a vulnerability is found (and they are always found), someone needs to fix it fast. **Features you don't build.** The features you skip in v1 become the features you wish you had in production. Multi-agent A2A messaging. VPC deployment options. SSO integration. These are all "v2 someday" until they're blockers. **Opportunity cost.** What could your team have shipped instead of building a control plane? ## The Honest Comparison | Factor | DIY Gateway | Sentrely | |---|---|---| | Time to first production agent | 2-3 months | Minutes | | Ongoing engineering maintenance | 20-40% of an engineer's time | Zero | | Feature completeness | Whatever you build | Maintained product | | SOC 2 / compliance evidence | You build the audit trail | Built-in | | Multi-agent support | Build it yourself | Included | | Slack / Telegram approvals | Build it yourself | Included | | VPC deployment option | Build it yourself | Enterprise tier | ## When DIY Makes Sense There are real cases where building your own is the right choice: - Your requirements are highly specific and no existing product fits - You have security requirements that prohibit third-party software - You're a platform company and the gateway is your core product - You have the engineering capacity and it's a strategic investment For most teams — especially those where building a control plane isn't core to your business — the math favors buying. You get to production faster, you spend less ongoing maintenance time, and you get features you'd never have prioritized building. The exception: if your gateway needs to be deeply customized for requirements that no managed product supports. Sentrely is open source — you can fork it if your needs genuinely require it. --- ### Sentrely vs Glean **URL:** https://sentrely.com/compare/vs-glean **Verdict:** Glean is enterprise AI search — answer questions across your knowledge base. Sentrely governs autonomous AI agents that take actions on your behalf. Different products, complementary stack. **Feature comparison:** | Feature | Sentrely | Glean | |---|---|---| | Cross-app enterprise search & RAG | ✕ | ✓ | | Per-agent RBAC for tool actions | ✓ | Search-result RBAC | | Audit trail of AI tool calls | ✓ | Search queries | | Human-in-the-loop approvals | ✓ | ✕ | | Multi-provider model routing | ✓ | ✕ | | 100+ data source connectors | ✕ | ✓ | | Built for autonomous code agents | ✓ | ✕ | | Permissions-aware retrieval | ✕ | ✓ | | Self-hosted / VPC option | ✓ | ✓ | **FAQ:** **Q: Is Sentrely a Glean alternative?** No. Glean is enterprise AI search — it indexes your Slack, Drive, GitHub, Confluence, Salesforce, etc. so employees can ask questions and get permission-aware answers. Sentrely is a control plane for autonomous AI agents that take actions on your behalf. Different problems entirely. If your question is 'where's the doc for X' — Glean. If your question is 'how do I keep my Claude agent from pushing to main without approval' — Sentrely. **Q: Can Glean replace what Sentrely does?** No. Glean retrieves and synthesizes information; it doesn't enforce policies on autonomous agents or audit their tool calls. Glean's agentic features (Glean Agents, Glean AI Actions) are designed for human-prompted workflows, not the autonomous, multi-step Claude/Cursor/Codex agents Sentrely governs. **Q: Can Sentrely replace what Glean does?** No. Sentrely doesn't index your enterprise content or build a knowledge base. If you need cross-app search and RAG over Slack/Drive/Confluence, Glean is purpose-built for that. Sentrely focuses on governing what AI agents do, not on what they know. **Q: Can I use Glean and Sentrely together?** Yes — increasingly common. Pattern: an autonomous agent running through Sentrely uses Glean as its knowledge retrieval layer. Sentrely audits and policy-checks every tool call (including Glean queries); Glean handles the permission-aware retrieval. Glean answers 'what does the codebase know'; Sentrely controls 'what the agent is allowed to do with that knowledge'. **Q: Pricing comparison?** Glean is enterprise contract — typical engagements start in the high five figures and scale with seat count and connector breadth. Sentrely is publicly priced ($199–$999/mo with custom Enterprise). They're not really competing line items because they sit in different parts of the stack. **Q: Which one do I need first?** Depends on your bottleneck. If your team can't find information across systems → Glean. If your AI agents are running unsupervised against production → Sentrely. If both, most teams start with Sentrely (lower price, immediate risk reduction) and add Glean later as the search/knowledge layer matures. **Article:** Glean and Sentrely live in the same broader category — "AI for the enterprise" — but they solve completely different problems and target different roles inside the buying organization. Glean answers a question your knowledge worker has. Sentrely governs an action an autonomous agent takes. One reads, the other writes; one serves humans, the other supervises AI. ## What Glean Is Glean is enterprise search and AI work assistant. It provides: - Indexed search across 100+ enterprise apps (Slack, Drive, GitHub, Confluence, Salesforce, Jira, Notion, etc.) - Permission-aware retrieval — users only see results they have access to - An AI assistant chat that synthesizes answers across that knowledge base - Glean Agents — task-specific AI assistants for things like onboarding, IT helpdesk, sales enablement - Workplace search analytics and content gap insights The typical Glean buyer is a 500+ person company struggling with knowledge fragmentation. The KPI is "time to answer" for employees. ## What Sentrely Is Sentrely is a control plane for autonomous AI agents — the Claude Code, Cursor, and Codex agents that engineering, ops, and growth teams are increasingly running in production. It provides: - Policy-based RBAC — what AWS actions, git repos, APIs each agent can touch - Full audit trail of every tool call, conversation, and approval - Human-in-the-loop gates — risky operations require Slack/Telegram approval - Multi-provider routing (Claude, OpenAI, Cursor, Codex) with automatic failover - Cost controls — per-agent token budgets, rate limits, spend alerts The typical Sentrely buyer is an engineering team that has adopted Claude Code or Cursor and needs operational controls before agents touch production. The KPI is "incidents prevented per quarter" and "time-to-audit-compliance". ## Where They Overlap (And Where They Don't) There's a thin slice of overlap: both involve AI, both touch enterprise data, both need RBAC. But the RBAC concerns are different: - **Glean RBAC** = "this user can see this Drive folder" - **Sentrely RBAC** = "this AI agent can call s3:PutObject on this bucket but not s3:DeleteObject" The first is about *information access*. The second is about *action authorization*. ## When You Need Both Increasingly common in 2026: an autonomous agent running through Sentrely calls Glean as its knowledge retrieval tool. The agent asks Glean "what's our refund policy" with a permission scope tied to the user who triggered the workflow; Glean returns the right answer; the agent then takes an action (issue refund, update CRM); Sentrely policy-checks and audits that action. In this stack: Glean is the brain (knowledge), Sentrely is the conscience (governance). They compose naturally. ## When You Don't Need Sentrely If your AI usage is bounded to Glean's chat and Glean Agents — humans asking questions, AI synthesizing answers from your indexed knowledge — Glean's built-in controls are likely sufficient. You only need Sentrely once your AI starts taking actions on systems Glean doesn't manage. ## When You Don't Need Glean If your AI agents already know what they need (because the prompts give them the context, or because they're working over a known codebase), you may not need enterprise search at all. Sentrely doesn't replace Glean's knowledge retrieval, but if knowledge retrieval isn't your bottleneck, you can skip it. ## The Bottom Line Glean is for knowledge — read-side AI that answers questions across your company's data. Sentrely is for agents — write-side AI that takes actions on your company's systems. Different categories, complementary purchases. --- ### Sentrely vs LangSmith **URL:** https://sentrely.com/compare/vs-langsmith **Verdict:** LangSmith tells you what happened. Sentrely tells you AND prevents the wrong things from happening. **Feature comparison:** | Feature | Sentrely | LangSmith | |---|---|---| | Trace LLM calls and chains | ✓ | ✓ | | Policy-based RBAC (deny-by-default) | ✓ | ✕ | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | ✕ | | Block actions in real time | ✓ | ✕ | | Multi-provider model routing & failover | ✓ | ✕ | | Cost / token budgets per agent | ✓ | Limited | | Prompt evaluation & A/B testing | ✕ | ✓ | | LangChain-native debugging | ✕ | ✓ | **FAQ:** **Q: Is Sentrely a LangSmith alternative?** Partially. LangSmith is best in class for LLM observability — tracing, prompt evaluation, debugging chains. Sentrely is a control plane — it enforces policies, gates risky actions on human approval, and produces compliance audit trails. They overlap on tracing but solve different jobs. Many teams use both: LangSmith for development debugging, Sentrely for production governance. **Q: Why would I need both LangSmith and Sentrely?** LangSmith helps you debug and improve your agents during development. Sentrely keeps them safe in production. LangSmith won't stop an agent from pushing to main; Sentrely will. Sentrely won't help you A/B test prompts; LangSmith will. **Q: Can Sentrely replace LangSmith for tracing?** For agent-level traces (every tool call, API request, conversation) — yes. For deep LangChain-specific debugging or prompt evaluation suites — no. If you only need to see what your agents did and search/replay sessions, Sentrely covers it. **Q: Does Sentrely work with LangChain agents?** Yes. Sentrely sits at the network layer — it doesn't care which framework your agents use. LangChain, Claude Code, custom Python, all of it routes through Sentrely the same way. **Q: What's Sentrely's pricing vs LangSmith?** Sentrely is $199/mo (Team) / $999/mo (Business). LangSmith Plus starts at $39/user/month with usage-based add-ons. Different cost structures — LangSmith bills per developer, Sentrely bills per team. **Article:** LangSmith is a genuinely excellent tool. If you're debugging LLM applications, understanding prompt behavior, or tracing complex chains, it's one of the best options available. The comparison to Sentrely isn't about which is better — they solve fundamentally different problems. ## What LangSmith Does Well LangSmith is an LLM observability and evaluation platform. Its strengths: - **Tracing:** Full visibility into LLM calls, inputs, outputs, latency, token usage - **Debugging:** See exactly what prompt was sent, what was returned, where in a chain things went wrong - **Evaluation:** Run datasets against your prompts, measure quality over time - **Playground:** Test prompt variations with real traces - **Dataset management:** Store and version test cases If your primary question is "why is my LLM application producing bad outputs?", LangSmith is built for that. ## What LangSmith Doesn't Do LangSmith observes your agents. It doesn't control them. - **No policy enforcement.** LangSmith doesn't intercept agent requests and check whether they're allowed. It records what happened after. - **No RBAC.** No concept of per-agent identity with scoped permissions. - **No approval gates.** When an agent wants to delete a production resource, LangSmith doesn't gate that operation — it logs it (after it happens). - **No cost controls.** No token budgets, no session termination, no runaway loop protection. - **No Slack/Telegram integration** for human-in-the-loop approvals. - **No audit trail for compliance.** LangSmith's traces are great for debugging, but they're not the immutable, structured audit logs that SOC 2 and HIPAA auditors are looking for. ## The Core Difference LangSmith answers: "What did my agents do, and was the output good?" Sentrely answers: "What are my agents allowed to do, are they doing it, and if something goes wrong, how do I stop it?" One is retrospective analysis. The other is real-time control. | Capability | LangSmith | Sentrely | |---|---|---| | LLM call tracing | Excellent | Basic (action-level) | | Prompt debugging | Excellent | Not the focus | | Output evaluation | Excellent | Not the focus | | Real-time policy enforcement | No | Yes | | Per-agent RBAC | No | Yes | | Approval gates | No | Yes | | Cost controls / token budgets | No | Yes | | SOC 2 / compliance audit trail | Partial | Yes | | Slack / Telegram approvals | No | Yes | ## Using Both These tools complement each other. Many teams use LangSmith for development and evaluation — debugging prompts, testing against datasets, improving output quality — and Sentrely for production — enforcing policies, logging compliance evidence, managing costs. If you're choosing between the two for a single production use case, the question is: what's your primary risk? If it's output quality (hallucinations, wrong answers, poor performance), LangSmith is more relevant. If it's operational risk (unauthorized access, runaway costs, compliance), Sentrely addresses it. For teams running Claude Code agents specifically — agents that take actions in the world, not just generate text — the control plane layer is where most of the risk lives. --- ### Sentrely vs LiteLLM **URL:** https://sentrely.com/compare/vs-litellm **Verdict:** LiteLLM routes LLM API calls across 100+ providers. Sentrely governs what Claude Code agents are allowed to do. One is infrastructure plumbing; the other is an agent control plane. **Feature comparison:** | Feature | Sentrely | LiteLLM | |---|---|---| | 100+ LLM providers via unified API | Major providers | ✓ | | Policy-based RBAC for agent actions | ✓ | ✕ | | Audit trail of every tool call | ✓ | LLM calls only | | Human-in-the-loop approvals | ✓ | ✕ | | Multi-provider failover & routing | ✓ | ✓ | | Token / cost budgets per agent | ✓ | ✓ | | Caching & rate limiting | ✓ | ✓ | | Open-source / self-hostable | ✓ | ✓ | | Compliance evidence (SOC2) | ✓ | ✕ | **FAQ:** **Q: Is Sentrely a LiteLLM alternative?** For pure LLM proxying (one unified API across providers, caching, rate limiting), LiteLLM is excellent. Sentrely is a layer above — it does proxying plus enforces what your agents are allowed to do (RBAC), logs every tool call (not just LLM calls), and gates risky operations on human approval. Use LiteLLM if you only need API routing. Use Sentrely if you need agent governance. **Q: Can Sentrely replace LiteLLM?** For most teams running Claude / OpenAI / Cursor agents, yes. Sentrely supports the major providers natively. For teams that need 100+ obscure providers (AWS Bedrock, Azure OpenAI, Vertex AI, Cohere, Replicate, Together, etc.), LiteLLM has wider coverage. **Q: Can I run Sentrely in front of LiteLLM?** Yes — common pattern. Use LiteLLM as the LLM router (Claude / GPT / etc.) and Sentrely as the agent governance layer in front of it. Sentrely catches every tool invocation; LiteLLM handles the model routing underneath. **Q: Both are open-source — what's the difference at the OSS level?** LiteLLM is a Python library/proxy server you self-host. Sentrely's gateway is similarly self-hostable. The managed Sentrely product adds the dashboard, audit retention, multi-tenancy, SSO, and 24/7 ops on top. **Q: Does Sentrely add latency vs LiteLLM?** Both add a few milliseconds of overhead. Sentrely's policy check is in-memory and adds <5ms per request. For agent workflows that take 1-5 seconds end-to-end, the difference is invisible. **Article:** ## What LiteLLM Is LiteLLM is an open-source LLM proxy that provides a unified OpenAI-compatible API in front of 100+ LLM providers: OpenAI, Anthropic, Azure, Google Vertex, AWS Bedrock, Cohere, Mistral, and dozens more. The core value proposition: write your code once against the OpenAI format, and LiteLLM routes it to whatever model and provider you want. Change providers without changing application code. LiteLLM also includes budget management at the team/key level, spend tracking, rate limiting, and basic audit logging. It's widely adopted by DevOps and platform teams managing LLM infrastructure across a multi-model organization. ## What Sentrely Is Sentrely is a managed control plane for Claude Code agents specifically. Claude Code agents aren't just making LLM API calls — they're reading git repositories, pushing code, accessing S3 buckets, calling external APIs, making database queries. Sentrely governs all of those actions: what each agent is allowed to access, which operations require human approval, how much each agent can spend, and what the immutable audit trail looks like for compliance. ## The Gap That Matters LiteLLM knows about LLM calls. Sentrely knows about agent actions. When `claude-deploy-01` wants to push a commit to `main`, LiteLLM has no concept of that operation — it only sees the Claude API call that preceded it. Sentrely intercepts the git push itself, checks the agent's policy (`git:push on main requires_approval: true`), routes the approval request to Slack, and logs the outcome. The control surface is completely different. | Capability | LiteLLM | Sentrely | |---|---|---| | Multi-provider LLM routing | Yes — 100+ providers | No — Claude Code focused | | OpenAI-compatible proxy | Yes | No | | Provider fallbacks | Yes | No | | Semantic caching | No (requires add-ons) | No | | Per-agent RBAC (git, AWS, APIs) | No | Yes | | Human approval gates | No | Yes — Slack + Telegram | | Agent action audit trail | LLM calls only | Every agent action | | SOC 2 / HIPAA compliance evidence | Not designed for this | Yes | | Session kill switch | No | Yes | | Token budgets per agent session | Key-level budgets | Per-session hard caps | | Claude Code specific | No | Yes | | Self-hosted | Yes (open-source) | Managed (or Enterprise VPC) | | Pricing | Free / Enterprise | Starter $49/mo | ## When LiteLLM Is the Right Choice - Your team uses multiple LLM providers and wants to abstract provider differences - You're on a platform team managing LLM access for many teams and models - You want self-hosted, open-source infrastructure you own and operate - Your primary concern is routing, caching, and spend visibility at the API call level ## When Sentrely Is the Right Choice - You're running Claude Code agents that touch production systems: git repos, AWS, databases, external APIs - Individual agents need different access scopes (deploy agent vs. review agent vs. data agent) - Certain agent operations — pushing to main, deleting data, sending emails — need human approval - You need to pass a compliance audit covering AI agent operations - You need to answer "what did agent X do last Tuesday at 2pm?" ## Can You Use Both? LiteLLM handles the LLM call layer. Sentrely handles the agent action layer. They don't overlap much because they're operating at different levels — LiteLLM sees Claude API requests, Sentrely sees agent tool calls against real systems. If you're running Claude Code agents, the relevant question isn't "which LLM should this call go to" — it's "should this agent be allowed to do this, and does a human need to approve it first?" That's a Sentrely problem, not a LiteLLM problem. --- ### Sentrely vs Lyzr **URL:** https://sentrely.com/compare/vs-lyzr **Verdict:** Lyzr builds enterprise AI agents for you. Sentrely controls the Claude Code agents you build yourself. **Feature comparison:** | Feature | Sentrely | Lyzr | |---|---|---| | Provider-agnostic (Claude, GPT, Cursor, Codex…) | ✓ | ✕ | | Bring-your-own-subscription (OAuth, no API key) | ✓ | ✕ | | Policy-based RBAC (YAML, deny-by-default) | ✓ | Limited | | Full audit trail of every tool call | ✓ | Partial | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | ✕ | | Multi-provider failover | ✓ | ✕ | | No credits — pay per usage | ✓ | ✕ | | Pre-built named agents (HR, sales, marketing) | ✕ | ✓ | | Self-hosted / VPC deployment | ✓ | ✓ | | Starting price | $199/mo | Enterprise contract | **FAQ:** **Q: Is Sentrely an alternative to Lyzr?** Sentrely and Lyzr both work in the AI agent space, but they solve different problems. Lyzr builds pre-packaged enterprise agents (HR, sales, marketing) for you. Sentrely is a control plane that governs the Claude, Cursor, and Codex agents your engineering team builds itself. If you're an engineering team writing your own agents, Sentrely is a direct alternative. If you're a business leader buying a finished agent for HR, Lyzr is the better fit. **Q: Why would an engineering team choose Sentrely over Lyzr?** Engineering teams pick Sentrely because they want to build agents, not buy them. Sentrely doesn't write agents — it controls what the agents you build can do. Lyzr's strength is professional services and pre-built solutions; that's overkill (and expensive) when you have a Claude Code or Cursor team already shipping. Sentrely starts at $199/mo with no enterprise contract. **Q: Does Sentrely support multiple AI providers like Lyzr?** Yes. Sentrely is provider-agnostic — connect Claude, OpenAI / Codex, Cursor, and (coming soon) Gemini, Mistral, and Llama. You can mix providers inside one workspace and define automatic failover so an outage at one provider doesn't stop your agents. **Q: Can I run Sentrely on-premise like Lyzr?** Yes — the Enterprise plan includes private VPC deployment. The managed plan ($199–$999/mo) runs on Sentrely's infrastructure, but the gateway can be self-hosted (open source) or deployed inside your VPC for compliance reasons. **Q: What does Sentrely cost vs Lyzr?** Sentrely: $199/mo (Team), $999/mo (Business), Custom (Enterprise) — all with transparent per-seat and per-agent caps, no credits. Lyzr is enterprise-contract-only — pricing isn't published, but typical engagements include implementation services and run into five-to-six-figure annual contracts. **Q: Can I use both Sentrely and Lyzr together?** Yes. If your company already uses Lyzr for, say, an HR agent, you can route Lyzr's API calls through Sentrely as the gateway to add audit logging and policy enforcement. The two aren't competitors at the integration level — they sit at different layers of the stack. **Article:** Lyzr and Sentrely are both in the enterprise AI agent space, but they serve fundamentally different buyers with different needs. Understanding the distinction helps you choose the right tool — or determine whether you need both. ## What Lyzr Is Lyzr is a full enterprise AI agent platform. It provides: - Pre-built named agents (Diane for HR, Jazon for sales, Skott for marketing, Amadeo for banking) - No-code/low-code agent builder with templates - Enterprise deployment infrastructure (managed or on-premise) - Industry-specific solutions (banking, insurance, healthcare) - Professional services and implementation support Lyzr is most valuable when you want Lyzr to build the AI agent solution for your business function. The typical Lyzr customer is a large enterprise buying a managed AI solution — comparable to buying enterprise software, not building it. ## What Sentrely Is Sentrely is a managed control plane for Claude Code agents. It provides: - Policy enforcement (RBAC) for agents you build - Audit trails and compliance evidence - Human-in-the-loop approvals via Slack/Telegram - Cost controls and token budgets - Session management and kill switches - Multi-agent orchestration support Sentrely is most valuable when you're already building with Claude Code and need operational infrastructure around your agents. The typical Sentrely customer is an engineering team that writes their own agents and needs a control plane to run them safely in production. ## The Core Difference | Dimension | Lyzr | Sentrely | |---|---|---| | What it provides | Pre-built AI agents | Control plane for your agents | | Target user | Business buyer, enterprise | Engineering team | | Model flexibility | Multi-model (GPT, Claude, Llama) | Claude Code focused | | Build vs. configure | Configure pre-built agents | Control agents you build | | Pricing model | Enterprise contract | Starter $49 / Pro $199 / Enterprise | | Time to first value | Weeks (implementation) | Minutes (point agent at gateway) | | Customization | Within platform constraints | Full control over your agents | ## Who Should Use Each **Lyzr:** Your organization wants AI agents for specific business functions (HR, sales, customer service) and prefers to buy a managed solution rather than build it. You have budget for enterprise software and want professional services support. **Sentrely:** You're already building with Claude Code, or you want to build your own agents and control exactly what they do. You have engineering resources and prefer building over buying. ## Not Really Competitors Most teams facing this choice aren't actually choosing between them — they're choosing a different approach entirely. Lyzr's buyer is thinking "we need AI in our HR workflow, who can implement it for us?" Sentrely's buyer is thinking "we're building Claude agents and need production-grade controls." If you're an engineering team building Claude Code agents, Sentrely is purpose-built for your use case. If you're a business leader evaluating AI automation for a specific department, Lyzr may be a better fit. --- ### Sentrely vs MuleSoft Agent Fabric **URL:** https://sentrely.com/compare/vs-mulesoft-agent-fabric **Verdict:** MuleSoft Agent Fabric is an enterprise agent control plane built around Salesforce, Agentforce, and the MuleSoft platform. Sentrely is the same idea for the rest of us — governing the Claude Code and Codex agents your team already runs, managed and from $199/mo, with no six-figure platform commitment. **Feature comparison:** | Feature | Sentrely | MuleSoft Agent Fabric | |---|---|---| | Agent control plane (register, govern, observe) | ✓ | ✓ | | Built for Claude Code / OpenAI Codex agents | ✓ | Agentforce-first | | Deny-by-default policy RBAC per agent | ✓ | ✓ | | Human-in-the-loop approvals (Slack / Telegram) | ✓ | Limited | | Immutable per-tool-call audit trail | ✓ | ✓ | | MCP / A2A governance | ✓ | ✓ | | Managed, set up in minutes | ✓ | ✕ | | Transparent pricing from $199/mo | ✓ | ✕ | | Self-hostable | ✓ | ✕ | | Provider-agnostic, no platform lock-in | ✓ | Salesforce ecosystem | **FAQ:** **Q: Is Sentrely a MuleSoft Agent Fabric alternative?** Yes, for teams that don't live inside the Salesforce/MuleSoft ecosystem. Agent Fabric is an enterprise platform — register, orchestrate, govern, and observe agents across a large org, governed through MuleSoft Flex Gateway and tied to Agentforce. Sentrely delivers the same core control plane (deny-by-default policy, audit, approvals, MCP governance) as a managed product you can turn on in minutes, aimed at the Claude Code and Codex agents engineering teams already run. **Q: How is the pricing different?** MuleSoft Agent Fabric is enterprise/quote-based and generally tied to an existing MuleSoft or Salesforce agreement — realistically a five- to six-figure commitment. Sentrely is transparent and starts at $199/mo, including the governance features (RBAC, audit, approvals) that enterprise vendors typically gate behind their top tier. **Q: Do they govern the same kind of agents?** Agent Fabric is built around Salesforce's Agentforce and enterprise agents, with broad protocol support (MCP, A2A). Sentrely is built first for autonomous coding/ops agents — Claude Code and OpenAI Codex — that touch real infrastructure (AWS, GitHub, databases), and it's provider-agnostic rather than anchored to one vendor's ecosystem. **Q: We're a small team running Claude Code in production. Which fits?** Sentrely. Agent Fabric is excellent if you're already standardized on MuleSoft/Salesforce and need enterprise-wide orchestration. If you're a developer or platform team that just needs governance on the agents you're running today — without adopting a whole integration platform — Sentrely is the lighter, faster, cheaper path. **Q: Does the existence of Agent Fabric mean this category is real?** Very much so. When Salesforce ships and advertises an agent control plane to 'stop agent sprawl and gain control,' it validates that governing autonomous agents is now a required layer — not a nice-to-have. The question is just whether you need an enterprise platform to get it, or a focused product that does it for the agents you actually run. **Article:** ## What MuleSoft Agent Fabric Is MuleSoft Agent Fabric is Salesforce's enterprise **agent control plane** — a single place to register, orchestrate, govern, and observe every AI agent and MCP endpoint across a large organization, regardless of where each agent was built. Governance is enforced through MuleSoft's **Flex Gateway** (applying security and compliance policy to MCP and Agent2Agent interactions), with an **AI Gateway** layer for centralized LLM token/cost visibility and an **Agent Visualizer** that maps how agents connect and perform. It's a powerful fit for enterprises already standardized on MuleSoft and Salesforce/Agentforce that need to wrangle agent sprawl across many teams and systems. ## What Sentrely Is Sentrely is a managed control plane built first for the autonomous **coding and ops agents** teams actually run today — Claude Code and OpenAI Codex. Those agents don't just call an LLM; they read git repos, push code, touch AWS, query databases, and call external tools. Sentrely governs every one of those actions: deny-by-default policy per agent, human approval on risky operations via Slack or Telegram, an immutable per-tool-call audit trail, and cost controls — with no infrastructure to run and pricing that starts at $199/mo. ## The Gap That Matters Both are control planes. The difference is **altitude and weight**. Agent Fabric is an enterprise *platform play*: it shines when you're orchestrating many agents across a Salesforce-centric organization and can absorb a MuleSoft adoption. That power comes with platform gravity — integration work, enterprise contracts, and an ecosystem to commit to. Sentrely is a *product*: point your agents' gateway URL at it and you have policy, approvals, and audit in minutes, scoped to the Claude Code / Codex agents you're running right now. No MuleSoft. No six-figure floor. No ecosystem lock-in. So the honest framing: **MuleSoft Agent Fabric governs the enterprise's Agentforce agents. Sentrely governs the Claude Code agents your team already runs** — without buying a platform to do it. If you're a developer or a small-to-mid engineering team that needs governance today, Sentrely is the lighter, faster, cheaper path. If you're a Fortune 500 standardizing agent orchestration across the whole company on Salesforce, Agent Fabric is built for exactly that. --- ### Sentrely vs n8n / Make / Zapier **URL:** https://sentrely.com/compare/vs-n8n **Verdict:** n8n automates workflows. Sentrely controls autonomous AI agents. These are different problems. **Feature comparison:** | Feature | Sentrely | n8n / Make / Zapier | |---|---|---| | Visual workflow builder | ✕ | ✓ | | Built for autonomous AI agents | ✓ | Limited (AI nodes) | | Per-agent RBAC & policy enforcement | ✓ | ✕ | | Audit trail of every tool call | ✓ | Workflow runs only | | Human-in-the-loop approvals | ✓ | Wait nodes | | Multi-provider model routing & failover | ✓ | ✕ | | 1000+ pre-built integrations | ✕ | ✓ | | Code-first / API-driven | ✓ | Hybrid | **FAQ:** **Q: Is Sentrely an n8n alternative?** No — they solve different problems. n8n (and Make, Zapier) are for deterministic workflows: 'when X happens, do Y.' Sentrely is for autonomous agents that decide for themselves what to do based on a goal. If you can flowchart the work in advance, use n8n. If the work requires reasoning, use an AI agent — and use Sentrely to control it. **Q: Can I use n8n and Sentrely together?** Yes. A common pattern: n8n triggers an AI agent on some event (Slack message, form submission, calendar event), the AI agent runs through Sentrely's gateway, and the result flows back into n8n for downstream non-AI steps. Sentrely doesn't replace your automation tooling — it governs the AI parts of it. **Q: Why not just use n8n's built-in AI nodes?** n8n's AI nodes work well for deterministic LLM calls inside a workflow (summarize this, classify that). They're not designed for autonomous agents that loop, plan multi-step actions, and need policy enforcement. The moment your AI step needs RBAC, an audit trail, or approval gating, n8n stops being enough. **Q: How does Sentrely's UI compare to n8n's drag-and-drop builder?** Sentrely is policy-first, not workflow-first. You write a YAML policy describing what each agent can do, and the agent makes its own decisions inside those constraints. There's a dashboard for monitoring, but the 'workflow' concept doesn't really apply — agents decide their own paths. **Q: What about Zapier or Make for AI?** Same answer. Zapier and Make are excellent for connecting SaaS apps deterministically. They're not built for autonomous agents that need to reason and decide. If your AI use case fits in a single 'AI node' in a Zap, you're fine. If it spans multiple steps and needs governance, you need Sentrely. **Article:** Traditional workflow automation tools — n8n, Make, Zapier — solve a real problem: connecting systems and automating repetitive processes. If you need to "when a form is submitted, create a CRM record and send a Slack message," these tools are excellent. They're deterministic, predictable, and easy to reason about. Claude agents solve a different problem: reasoning about complex, unstructured situations and taking adaptive action. They're not automating a defined workflow — they're applying judgment to situations that can't be fully specified in advance. ## The Fundamental Difference **n8n/Make/Zapier:** You define the workflow. A trigger fires. Each node executes a predetermined action. The output is predictable because the path is fixed. **Claude agents:** You describe a goal. The agent figures out the steps. The path is chosen by the agent based on the situation it encounters. The output depends on judgment, not just execution. This difference matters for control. You don't need to "approve" an n8n node executing — you defined exactly what it would do. You might need to approve a Claude agent pushing to your main branch — because the agent's judgment about when it's done isn't the same as yours. ## What n8n/Make/Zapier Do Well - **Deterministic automation.** You know exactly what will happen before it runs. - **400+ integrations.** If it has an API, it probably has a node. - **Visual workflow design.** Non-engineers can build and understand workflows. - **Reliable execution.** Runs the same way every time. - **Lower cost for simple automation.** Hard to beat for connect-A-to-B workflows. ## What They Don't Do - **Adaptive reasoning.** Can't handle situations the workflow designer didn't anticipate. - **Natural language task understanding.** Can't read a code review request and figure out what to check. - **Context-aware decision making.** Can't assess whether a refactoring is safe based on the broader codebase. - **Control for autonomous agents.** No concept of per-agent identity, policy scoping, or approval gates for AI-driven actions. ## The Decision Framework | If you need... | Use | |---|---| | "When X happens, do Y" automation | n8n / Make / Zapier | | "Figure out how to do X" autonomous agents | Claude Code + Sentrely | | Both in the same stack | Both — they complement each other | ## Where They Overlap Some teams use n8n to trigger Claude agents — an n8n workflow receives a webhook, extracts relevant data, and passes it to a Claude agent to handle the reasoning-heavy part. The n8n handles the plumbing; the Claude agent handles the judgment. In this architecture, Sentrely controls the Claude agent portion while n8n handles the deterministic orchestration. ## The Control Question The reason you need a control plane for Claude agents and not for n8n workflows is exactly this: n8n does what you told it to do. Claude agents do what they think is right. When the question is "did the workflow execute correctly?", n8n's own logs answer it. When the question is "did the agent make a good decision about what to do?", you need a control plane that enforces policies, requires approvals, and gives you a kill switch. The autonomy that makes Claude agents more powerful than n8n is exactly what requires more sophisticated control. --- ### Sentrely vs Raw Claude API (no gateway) **URL:** https://sentrely.com/compare/vs-no-gateway **Verdict:** Raw API access is fine for local dev and prototypes. It's not acceptable in production. **Feature comparison:** | Feature | Sentrely | Raw Claude API (no gateway) | |---|---|---| | Policy-based RBAC (per-agent permissions) | ✓ | ✕ | | Audit trail of every tool call | ✓ | ✕ | | Human-in-the-loop approvals | ✓ | ✕ | | Cost / token budgets and alerts | ✓ | ✕ | | Multi-provider failover | ✓ | ✕ | | Compliance evidence (SOC2, HIPAA) | ✓ | ✕ | | Direct provider API access | ✓ | ✓ | | Setup time | 10 minutes | Instant | **FAQ:** **Q: Why can't I just use the raw Claude API?** You can — for prototypes, internal tools, and local dev. The problem starts in production: a raw API integration has no policy layer, no audit log, no rate limits, and no kill switch. When (not if) an agent goes rogue, you have no way to stop or investigate it. A control plane like Sentrely is what turns 'I tried Claude' into 'I run Claude in production.' **Q: What goes wrong without a gateway?** Three common incidents: (1) An agent hits an infinite loop and burns through your monthly token budget in 20 minutes. (2) An agent gets confused and pushes broken code to main without human review. (3) Compliance audit asks 'show me everything Claude touched on March 4' — and you can't, because nothing was logged. **Q: Isn't Anthropic's API rate-limited already?** Yes, at the org level — but that's a sledgehammer. You can't say 'this agent gets 100k tokens/day, that one gets 10k' or 'this agent can only call S3 read, not write' through Anthropic's rate limits. Sentrely gives you per-agent, per-resource granularity. **Q: How much overhead does Sentrely add?** Less than 5ms per request. The gateway runs policy checks in memory and forwards to the upstream provider. For a typical Claude conversation that takes 1-3 seconds end-to-end, the gateway is invisible. **Q: Can I migrate from raw API to Sentrely without code changes?** Mostly yes. Set the `ANTHROPIC_BASE_URL` env var to your Sentrely gateway endpoint instead of `api.anthropic.com`. Your existing code keeps working; Sentrely now enforces policies, logs everything, and routes risky calls to approval queues. **Article:** The simplest way to run a Claude agent is also the most common way to end up with an incident: give the agent your credentials and let it go. This isn't a criticism of Anthropic's API. It's an excellent API. The problem is what it doesn't provide: the operational layer you need to run agents safely against real systems. ## What You Get With Raw API Access When you point Claude Code directly at external services — no gateway in between — you get: - **Full access, whatever your credentials allow.** The agent can do anything you can do. If your AWS key has admin access, the agent has admin access. - **No audit trail.** The Anthropic API logs your token usage. It doesn't log what your agent did with those tokens, which systems it touched, or what it changed. - **No policy enforcement.** There's no layer between the agent and the resources it can reach. If you give it a command and it decides to take a broad interpretation, there's nothing to stop it. - **No cost controls.** You'll know how much you spent on your next invoice. You won't know until it arrives. - **No approval gates.** Destructive operations run if the agent decides to run them. - **No agent identity.** If you have multiple agents sharing credentials, your audit trail is useless. This is acceptable for: local development, prototypes, demos, personal tools with limited blast radius. ## What Sentrely Adds | Capability | Raw API | Sentrely | |---|---|---| | Audit trail | None | Every action, immutable | | RBAC / policy enforcement | None | Per-agent YAML policies | | Human approval gates | None | Slack / Telegram / dashboard | | Cost controls | Invoice after the fact | Per-session budgets + alerts | | Agent identity | Shared credentials | Per-agent identity | | Runaway loop protection | None | Circuit breaker + token limits | | Kill switch | Kill the process | Session terminate via dashboard | | Compliance evidence | None | Structured, queryable audit log | ## The Decision Framework **Use raw API access when:** - You're building a prototype or proof of concept - The agent only has access to your local machine - No production data, no production credentials, no production systems - You're willing to lose anything the agent might touch **You need a control plane when:** - Any agent touches production systems - Multiple agents share an environment - You have a compliance requirement (SOC 2, HIPAA, GDPR) - You're running agents overnight or without human supervision - Token costs matter to your budget The gap between "this works in my terminal" and "this is safe to run against production" is exactly the gap a control plane fills. Sentrely adds a layer between your agents and the world — a layer that enforces policies, logs everything, and keeps humans in control of the decisions that matter. --- ### Sentrely vs Paperclip **URL:** https://sentrely.com/compare/vs-paperclip **Verdict:** Paperclip models your agent fleet as a company org chart — roles, budgets, reporting lines. Sentrely models it as a policy-enforced control plane — RBAC, audit trails, approval gates. Both govern agents; the metaphors and target use cases differ. **Feature comparison:** | Feature | Sentrely | Paperclip | |---|---|---| | Policy-based RBAC (YAML, deny-by-default) | ✓ | Role-based | | Audit trail of every tool call | ✓ | ✓ | | Human-in-the-loop approvals | Slack/Telegram | Org-chart routing | | Multi-provider model routing & failover | ✓ | Limited | | Code-first config (YAML) | ✓ | ✕ | | Org-chart / role hierarchy modeling | ✕ | ✓ | | Budget & reporting line metaphors | ✕ | ✓ | | Engineering-team focus | ✓ | Business-team focus | **FAQ:** **Q: What's the core difference between Sentrely and Paperclip?** Both govern AI agents, but with different metaphors. Paperclip models your fleet like a company org chart — agents have roles, budgets, and reporting lines. Sentrely models it as a policy engine — every action checks against a YAML policy, with explicit allow/deny decisions. If you think in org charts, Paperclip. If you think in IAM policies, Sentrely. **Q: Which is better for engineering teams?** Sentrely. Engineers tend to think in code-first config (YAML, versioned in git, reviewed in PRs). Paperclip's org-chart abstractions are more natural for business-side governance. Sentrely's deny-by-default + explicit allow lists matches how engineers already think about IAM and Kubernetes RBAC. **Q: Which is better for business / ops teams?** Paperclip is built around metaphors business teams already use — roles, budgets, approval chains. If your governance committee thinks in org charts, Paperclip's UI will resonate. **Q: Can I migrate between them?** Yes — both export audit logs in standard formats and use HTTP-level interception. Migrating policies between the two requires manual translation (org-chart roles to YAML policies, or vice versa) but the agent integration layer is similar. **Q: Pricing comparison?** Sentrely is publicly priced ($199–$999/mo + Enterprise custom). Paperclip pricing varies by deployment model and isn't fully transparent at the time of writing. For typical engineering teams, Sentrely's transparent pricing and self-serve onboarding tend to be the easier choice. **Article:** ## What Paperclip Is Paperclip is an open-source "human control plane for AI labor." Its core metaphor is the org chart: you create an agent organization with roles, reporting structures, and goals — and the agents work within that structure. The board (you) can hire agents, approve strategic directions, override decisions, and terminate agents. Each agent has a budget that auto-pauses when hit. Paperclip is MIT-licensed, self-hosted, and runtime-agnostic — it works with Claude Code, Cursor, or any agent runtime. It's an impressive project with genuine thinking about the human oversight problem. ## What Sentrely Is Sentrely is a managed control plane for Claude Code agents, focused on policy enforcement, audit trails, and production-grade governance. It models agent permissions as YAML policies: `claude-deploy-01` can push to `feature/*` but not `main`, can read from specific S3 prefixes, requires human approval for production deployments. Human oversight happens through real-time Slack and Telegram approval gates. Every action produces an immutable audit log designed for compliance auditors. Sentrely is a fully managed service — no infrastructure to operate. ## Same Problem, Different Approach Both products address the same core problem: autonomous AI agents need human oversight and bounded permissions. The approaches diverge significantly. **Paperclip's approach:** Model agent governance as organizational structure. Agents have roles in a hierarchy, goals that cascade from the top, and budgets allocated like headcount. The metaphor makes sense for teams thinking about AI labor as something that needs management like human labor. **Sentrely's approach:** Model agent governance as infrastructure policy. Agents have RBAC policies scoped to specific resources, approval gates for specific operations, and immutable audit trails. The metaphor comes from security engineering — think AWS IAM, not an org chart. | Dimension | Paperclip | Sentrely | |---|---|---| | Core metaphor | Company org chart | Infrastructure policy engine | | Licensing | MIT open-source | Managed service (OSS also available) | | Self-hosted | Yes (required) | Managed cloud or Enterprise VPC | | Runtime support | Claude Code, Cursor, others | Claude Code focused | | Per-agent budgets | Yes | Yes — per-session hard caps | | Human approval gates | Board-level approvals | Slack/Telegram with full context | | RBAC (resource-level) | No — role/goal based | Yes — per-resource, per-action | | Immutable audit trail | Basic | Purpose-built for compliance | | Slack/Telegram integration | No | Yes — Butler bot | | Web dashboard | No | Yes | | Multi-agent A2A messaging | No | Yes | | SOC 2 / HIPAA evidence | No | Yes | | Pricing | Free (self-hosted) | Starter $49/mo | ## When Paperclip Makes Sense Paperclip is a strong fit if you: - Want fully self-hosted, open-source agent governance with no vendor dependency - Prefer the organizational metaphor — thinking about AI labor like managing a team - Are using multiple agent runtimes beyond Claude Code - Want to experiment with agent governance without committing to a paid service ## When Sentrely Makes Sense Sentrely is a stronger fit if you: - Need resource-level RBAC: "this agent can read S3 bucket A but not B, can push to feature branches but not main" - Need human-in-the-loop approval workflows that route to Slack with full context - Need compliance evidence for SOC 2 or HIPAA — immutable, auditor-ready logs - Don't want to operate infrastructure yourself - Are deeply invested in the Claude Code ecosystem ## The Honest Assessment Paperclip is the most philosophically interesting open-source project in this space. If you're willing to run infrastructure and want maximum control over the governance model, it's worth exploring. Sentrely trades the openness of DIY for the completeness of a managed product: resource-level RBAC, Slack approval workflows, immutable audit trails, and a dashboard that works out of the box. For teams that need to show SOC 2 or HIPAA evidence — or who just don't want to build and operate their own governance infrastructure — Sentrely closes the gap faster. The real competition between the two isn't features. It's the build-vs-buy question applied to agent governance: do you want to own it or operate it? --- ### Sentrely vs Portkey **URL:** https://sentrely.com/compare/vs-portkey **Verdict:** Portkey is the right gateway if you need 50+ LLM providers, semantic caching, and prompt management. Sentrely is right if you need per-agent RBAC, human approval gates, and compliance-grade audit trails for Claude Code agents. **Feature comparison:** | Feature | Sentrely | Portkey | |---|---|---| | 50+ LLM providers via unified API | Major providers | ✓ | | Per-agent RBAC (AWS, git, tool-level) | ✓ | ✕ | | Audit trail of every tool call | ✓ | LLM calls only | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | ✕ | | Multi-provider failover & routing | ✓ | ✓ | | Semantic caching | ✕ | ✓ | | Prompt versioning & management | ✕ | ✓ | | Compliance evidence (SOC2, HIPAA) | ✓ | ✓ | | Open-source / self-hostable | ✓ | ✓ | **FAQ:** **Q: Is Sentrely a Portkey alternative?** For pure LLM gateway use cases (multi-provider routing, semantic caching, prompt management), Portkey is purpose-built. Sentrely is a control plane focused on agent governance — RBAC for agent actions (not just LLM calls), human approval gating, and tool-level audit. Pick Portkey for LLM ops; pick Sentrely for agent governance. **Q: Can I use Portkey and Sentrely together?** Yes — common pattern. Portkey handles LLM provider routing, caching, and prompt management at the model layer. Sentrely sits in front of your agents and enforces what they're allowed to do at the tool/resource layer. Many teams use both for production deployments. **Q: What about semantic caching and prompt management?** Sentrely doesn't do these — they're Portkey's strength. If those features matter (e.g. you're optimizing latency on retrieval-heavy workflows), Portkey is the right primary tool. If your priority is governance and approvals, Sentrely is. **Q: Pricing comparison?** Portkey has a generous free tier (with usage limits) and paid tiers based on volume. Sentrely is $199/mo (Team) / $999/mo (Business) with transparent per-agent and per-seat caps. Different cost models — Portkey scales with LLM call volume, Sentrely with team size. **Q: Does Sentrely have prompt versioning?** Not natively. We track every conversation and tool invocation in the audit log, but for collaborative prompt iteration and A/B testing, Portkey or LangSmith are better tools. **Article:** ## What Portkey Is Portkey is a full-featured AI gateway and observability platform. It sits in front of your LLM calls and provides routing, caching, guardrails, rate limiting, and cost tracking across more than 50 LLM providers — OpenAI, Anthropic, Azure, Bedrock, Cohere, and dozens more. Teams that use multiple models, or want to swap providers without changing application code, get real value from it. Portkey's strongest features are provider flexibility and developer experience: three-line integration, semantic caching that reduces duplicate LLM calls, and a prompt management studio where non-engineers can update prompts without code deploys. ## What's Different The fundamental difference is orientation: Portkey is **provider-centric** (routing across many models) while Sentrely is **agent-centric** (governing what Claude Code agents can do in production). This gap shows up concretely in how each product handles the problems that matter most when running autonomous agents: **Policy enforcement.** Portkey's access controls operate at the API key and workspace level — which users can call which models. Sentrely's RBAC operates at the per-agent, per-resource level: `claude-deploy-01` can push to `feature/*` branches but not `main`, can read from `s3://acme/reports/*` but not write, can call Stripe's read API but not `charges:create`. These are fundamentally different control surfaces. **Human-in-the-loop approvals.** When a Claude agent wants to push to a production branch, delete a database record, or send an email to 10,000 customers, you need a human to approve it before it executes. Sentrely has built-in Slack and Telegram approval workflows for exactly this. Portkey has no equivalent — it doesn't gate individual agent operations on human review. **Agent identity.** Sentrely assigns each Claude agent its own identity with its own credential and policy. Every audit log entry tells you which specific agent did what. Portkey tracks by virtual key, which is closer to team-level attribution than per-agent attribution. **Compliance audit trail.** Sentrely's audit logs are designed specifically for SOC 2 and HIPAA auditors — immutable, structured, exportable, with field-level detail auditors require. Portkey's logs are primarily for debugging and cost optimization, not external audit evidence. ## The Comparison Table | Capability | Portkey | Sentrely | |---|---|---| | LLM provider support | 50+ (OpenAI, Claude, Gemini, Mistral, etc.) | Claude Code focused | | Semantic caching | Yes — cuts duplicate call costs | No | | Prompt management studio | Yes — non-engineers update prompts | No | | Per-agent RBAC | No — key/workspace level only | Yes — per-agent, per-resource policies | | Human approval gates | No | Yes — Slack + Telegram | | Immutable audit trail | No — logs for debugging/cost | Yes — designed for compliance auditors | | Claude Code agent integration | Generic LLM gateway | Purpose-built for Claude Code | | Token budgets (per session) | Rate limiting + spend tracking | Hard per-session budget caps | | Session kill switch | No | Yes | | A2A messaging | No | Yes | | Multi-agent orchestration | No | Yes | | Pricing | From $59.99/mo | Starter $49/mo | ## When Portkey Is the Right Choice - Your team uses multiple LLM providers and wants a single gateway for all of them - You want to swap from OpenAI to Claude without changing application code - Semantic caching matters for your use case (high-volume repetitive queries) - A prompt management studio for non-technical stakeholders is a priority - Your control needs are at the team/workspace level, not per-agent ## When Sentrely Is the Right Choice - You're running Claude Code agents in production against real infrastructure - Individual agents need different permission scopes (deployment agent vs. code review agent) - Certain agent operations require human sign-off before executing - You need to pass a SOC 2 or HIPAA audit that covers AI agent operations - You need to prove to an auditor exactly which agent did what and when ## Can You Use Both? In theory — Portkey as the LLM routing layer, Sentrely as the agent governance layer. In practice, the overlap is substantial enough that most teams choose one. If Claude is your primary model and agent governance is your priority, Sentrely handles both. If multi-provider flexibility is the priority and governance is secondary, Portkey is the stronger starting point. The underlying question: are you solving a provider routing problem or an agent governance problem? They look similar from the outside but lead to very different product needs. --- ### Sentrely vs Retool **URL:** https://sentrely.com/compare/vs-retool **Verdict:** Retool builds internal tools your team clicks through. Sentrely governs the autonomous AI agents acting on your behalf. Different layers — many teams use both. **Feature comparison:** | Feature | Sentrely | Retool | |---|---|---| | Drag-and-drop UI for internal tools | ✕ | ✓ | | Per-agent RBAC for AWS, git, tools | ✓ | User-level | | Audit trail of every AI tool call | ✓ | User actions | | Human-in-the-loop approvals (Slack/Telegram) | ✓ | Approval workflows | | Multi-provider AI model routing | ✓ | Retool AI | | Built for autonomous AI agents | ✓ | Limited (AI Actions) | | 1000+ data source connectors | ✕ | ✓ | | Code-first config (YAML) | ✓ | Hybrid | | VPC / self-hosted option | ✓ | ✓ | **FAQ:** **Q: Is Sentrely a Retool alternative?** No — Retool is for building internal admin tools that humans use (CRUD UIs, dashboards, support ops panels). Sentrely is for governing autonomous AI agents that act on your behalf without a UI. They sit at different layers of the stack and don't compete directly. **Q: What about Retool AI / AI Actions?** Retool AI lets you embed LLM calls inside your internal tool workflows — useful for things like 'summarize this ticket' inside a support panel. It's not designed for autonomous, multi-step Claude or Cursor agents that need policy enforcement, audit trails, and approval gating. Sentrely is purpose-built for that. **Q: Can I use Retool and Sentrely together?** Yes — common pattern. Build an internal Retool app that lets your ops team trigger AI agents on demand. The agents themselves run through Sentrely's gateway, so every AI action is policy-checked and audited. You get Retool's UI, Sentrely's governance. **Q: Why does my team need Sentrely if Retool already has user-level RBAC?** Retool's RBAC governs what humans can click. Sentrely's RBAC governs what AI agents can do — separate concern. An autonomous agent running in your CI pipeline isn't a Retool user; it doesn't have a Retool session. Sentrely sits at a different layer to control non-human actors. **Q: Pricing comparison?** Retool: $10/standard user/month, $50/business-user/month, with usage-based add-ons. Sentrely: $199/mo (Team) / $999/mo (Business) — per-team pricing rather than per-user. They're not really competing line items. **Article:** Retool and Sentrely solve different problems for different actors. Retool is for the humans on your team — your ops, support, finance, and growth folks who need internal tools to do their jobs. Sentrely is for the AI agents on your team — the autonomous Claude, Cursor, and Codex agents writing code, sending invoices, and managing cloud resources without a human at the keyboard. If your "internal tooling" stack is humans clicking buttons in admin panels, you need Retool. If it's autonomous agents acting on your behalf, you need Sentrely. If it's both — and increasingly, it is — you need both. ## What Retool Is Retool is a drag-and-drop builder for internal tools. It provides: - A visual canvas with 100+ pre-built UI components (tables, forms, charts) - Connectors to 1000+ data sources (Postgres, Salesforce, Stripe, REST APIs) - User-level RBAC and audit logs of human actions - Approval workflows for sensitive operations - Recently: Retool AI / AI Actions for embedding LLM calls inside tools The typical Retool customer is a 20-500-person company that needs ops tools faster than their engineering team can build from scratch. Customer support panels, refund processors, lead qualification tools — anything where a human needs a quick UI over data. ## What Sentrely Is Sentrely is a control plane for autonomous AI agents. It provides: - Policy-based RBAC — what AWS actions, git repos, APIs each agent can touch - Full audit trail of every tool call, conversation, and API request - Human-in-the-loop approvals via Slack/Telegram - Multi-provider routing (Claude, OpenAI, Cursor, Codex) with automatic failover - Cost controls — per-project token budgets, rate limits, spend alerts The typical Sentrely customer is an engineering team running Claude Code, Cursor, or Codex agents in production and needs operational controls. CI/CD agents, customer-support bots, billing automations, code-review agents — anything where AI is making decisions on your behalf. ## When You Need Both The classic 2026 pattern: a Retool internal panel for your ops team, with buttons that trigger AI agents to actually do the work. Click "Process refund queue" in Retool → Sentrely-governed billing-agent reads queue, validates each refund against policy, escalates over-threshold ones to a human in Slack, processes the rest, writes audit log. Retool gives your ops team the front door. Sentrely makes sure the AI agents behind that door don't go rogue. ## When You Don't Need Sentrely If your "AI" usage is purely LLM calls inside Retool workflows ("summarize this ticket"), Retool AI alone is probably enough. The moment you have autonomous agents running outside Retool — in your CI pipeline, on a server, in a Lambda — you need a control plane, and Sentrely is purpose-built for that role. ## When You Don't Need Retool If your internal stack is API-first and your ops team is technical (engineers, SREs), you might not need a UI builder at all. Sentrely's dashboard plus your existing CLIs and Slack integrations can replace many internal admin panels — especially when the workflows are AI-driven. ## The Bottom Line Retool builds the surfaces humans interact with. Sentrely controls the agents acting in the background. Different problems, different tools, complementary stack. --- ### Sentrely vs TrueFoundry **URL:** https://sentrely.com/compare/vs-truefoundry **Verdict:** TrueFoundry governs AI systems at the environment and team level — great for platform teams managing LLM infrastructure. Sentrely governs at the per-agent, per-resource level — designed for teams running autonomous Claude agents with real production access. **Feature comparison:** | Feature | Sentrely | TrueFoundry | |---|---|---| | Per-agent (not per-team) RBAC | ✓ | ✕ | | Per-resource action policies (AWS/git/tools) | ✓ | Limited | | Human-in-the-loop approvals | ✓ | ✕ | | Audit trail of every tool call | ✓ | LLM calls only | | Multi-provider model routing & failover | ✓ | ✓ | | Self-hosted ML platform | ✕ | ✓ | | Model serving / inference deployment | ✕ | ✓ | | VPC / on-prem deployment | ✓ | ✓ | **FAQ:** **Q: Is Sentrely a TrueFoundry alternative?** For pure agent governance, yes. TrueFoundry is a broader ML platform that includes an LLM gateway, model serving, and infrastructure orchestration. Sentrely is purpose-built for the agent governance slice — RBAC, approvals, and audit at the per-agent level. If you need a full ML platform, TrueFoundry. If you're focused on Claude/Cursor/Codex agents, Sentrely. **Q: Why pick Sentrely over TrueFoundry's gateway?** Granularity. TrueFoundry governs at the team/environment level — good for platform engineering. Sentrely governs at the per-agent, per-resource level — necessary when you have autonomous agents touching production. Different problems, different tools. **Q: Can I use both?** Yes. Many enterprises use TrueFoundry for ML/model infrastructure and Sentrely as the agent-specific governance layer. They sit at different parts of the stack. **Q: Does Sentrely deploy in our VPC like TrueFoundry?** Yes — Enterprise plan includes VPC deployment. The gateway can run in your AWS/Azure/GCP account, with audit logs streaming to your own S3. **Q: What about model serving and inference deployment?** Out of scope for Sentrely. We're a gateway, not a model serving platform. If you need to host your own models, TrueFoundry, BentoML, or KServe are better fits. Sentrely sits in front of whichever inference layer you choose. **Article:** ## What TrueFoundry Is TrueFoundry is an enterprise AI platform with an AI Gateway component that covers LLM, MCP, and agent traffic under a single control plane. It governs at the environment level: which teams can use which models, cost controls per team or project, observability across your AI stack. TrueFoundry also provides the surrounding deployment platform — model hosting, fine-tuning, vector databases — making it a broader MLOps/LLMOps offering than a pure gateway. ## What Sentrely Is Sentrely is a managed control plane built specifically for Claude Code agents. It governs at the per-agent, per-resource level: `claude-infra-01` can invoke specific ECS services but not modify IAM roles; `claude-research-01` can read from approved S3 prefixes but not write; any push to `main` triggers a Slack approval gate before execution. Every action is logged to an immutable audit trail built for compliance auditors. ## The Governance Level That Matters The key architectural difference is where governance is applied. **TrueFoundry governs at the environment/team level.** "The data science team can use GPT-4 and Claude 3.5 Sonnet. Their monthly budget is $5,000. Requests are rate-limited to 100 RPM." This is infrastructure governance — appropriate for a platform team managing LLM access across an organization. **Sentrely governs at the per-agent/per-resource level.** "This specific agent can push to feature branches, read from this S3 prefix, and call these Stripe endpoints — but any production deployment requires human approval." This is operational governance — appropriate for teams running agents with real production access. These aren't competing — they're different layers. But if your problem is "my Claude Code agents have too much access and I need approval workflows for high-risk operations," TrueFoundry doesn't solve that. Sentrely does. | Capability | TrueFoundry | Sentrely | |---|---|---| | Governance granularity | Team/environment level | Per-agent, per-resource level | | Multi-provider LLM support | 250+ models | Claude Code focused | | Model hosting + fine-tuning | Yes — full MLOps platform | No | | MCP gateway | Yes | Integrated | | Per-agent RBAC (git, AWS, APIs) | No | Yes | | Human approval gates | No | Yes — Slack + Telegram | | Immutable compliance audit trail | Observability focus | SOC 2 / HIPAA designed | | Claude Code specific | No | Yes | | Self-hosted / on-premises | Yes | Enterprise VPC tier | | Managed cloud | Yes | Yes | | Pricing | Enterprise (not listed) | Starter $49/mo | ## When TrueFoundry Makes Sense TrueFoundry is the right choice when: - You need a full MLOps platform (model hosting, fine-tuning, vector DB) in addition to a gateway - Your primary concern is LLM access management across multiple teams and models - You want on-premises deployment with enterprise support contracts - Environment-level governance (team budgets, model allowlists) covers your compliance needs ## When Sentrely Makes Sense Sentrely is the right choice when: - Claude Code agents are accessing real production infrastructure — git repos, AWS, APIs, databases - Individual agents need different permission scopes that aren't captured by team-level access control - Specific operations (production deploys, data deletion, customer communications) need human approval before execution - You need per-agent audit evidence that satisfies SOC 2 or HIPAA requirements - You don't need the broader MLOps platform, just the governance layer ## The Scope Question TrueFoundry is a platform that includes a gateway. Sentrely is a gateway built for a specific, hard use case: autonomous agents with real production access that need tight governance. If you need the full MLOps stack and governance is one of several requirements, TrueFoundry is worth evaluating. If your problem is specifically "Claude Code agents in production with controlled access, approval gates, and compliance trails," Sentrely is purpose-built for exactly that. --- ### Sentrely vs Viktor **URL:** https://sentrely.com/compare/vs-viktor **Verdict:** Viktor does the work. Sentrely makes sure the work stays safe. **Article:** Viktor is a $75M-funded AI coworker that lives in Slack and Teams. It connects to 3,000+ tools, executes tasks, builds dashboards, writes code, and automates workflows. Sentrely is the control plane that governs AI agents — enforcing policies, logging every action, and gating risky operations on human approval. They solve different problems. Viktor *is* the agent. Sentrely sits *around* your agents. ## Viktor's Strengths Viktor is impressive as an execution engine: - **Lives in Slack/Teams.** Message it like a coworker. No new UI to learn. - **3,000+ integrations.** Stripe, HubSpot, Google Ads, GitHub, Linear — connects via OAuth in seconds. - **Builds things.** Creates dashboards, internal tools, reports, PDFs. Tangible output. - **Persistent memory.** Learns your company's processes over time ("Skills" system). - **Zero setup.** No infrastructure, no API keys to configure, no code. - **Scales fast.** 2,000+ organizations, $15M ARR within weeks of launch. If you need an AI that does work for you — pulls reports, manages campaigns, researches leads — Viktor is excellent at that. ## What Viktor Doesn't Do Viktor is the agent, not the governance layer. It doesn't solve: - **Policy enforcement.** No mechanism to define what Viktor *can't* do. It has access to everything you connect. - **Per-agent RBAC.** One Viktor instance per workspace. No scoping different agents to different permission sets. - **Approval gates for risky operations.** Viktor executes immediately. No "pause and ask a human" before destructive actions. - **Immutable audit trail.** Viktor logs exist, but they're not designed for SOC 2 evidence or compliance audits. - **Multi-agent orchestration.** Viktor is one agent. If you're running multiple agents (Claude, Codex, custom), Viktor doesn't govern the fleet. - **Cost controls.** Credit-based pricing, but no per-agent token budgets or automatic session termination when spending spikes. - **True isolation.** Viktor runs in their cloud. No microVM isolation between your tasks, no way to deploy in your VPC. ## The Fundamental Difference | | Viktor | Sentrely | |---|---|---| | **What it is** | The agent itself | The control plane around agents | | **Core job** | Execute tasks | Enforce policies on tasks | | **Interface** | Slack/Teams native | Dashboard + Slack/Telegram approvals | | **Integrations** | 3,000+ (does the work) | 100+ (gates the access) | | **Security model** | SOC 2, SSO (enterprise) | RBAC, policy YAML, microVM, audit trail (all plans) | | **Pricing** | $50–$50k/mo (credits) | $199–$999/mo (per-agent, flat) | | **Target buyer** | Ops teams wanting automation | Eng teams governing their agents | ## When to Use Viktor - You want an AI to *do* operational work (reports, automations, data pulls) - You're a non-technical team that needs automation without code - Your risk tolerance is high — you trust the agent with broad access - You want one general-purpose agent, not a fleet of specialized ones ## When to Use Sentrely - You already have agents (Claude, Codex, custom) and need governance - You need to scope what each agent can access (this agent: read-only S3; that agent: full deploy) - Compliance requires an immutable audit trail of every agent action - You want approval gates before agents touch production - You're running multiple agents and need fleet-wide visibility - You want agents in your VPC, not a vendor's cloud ## Can You Use Both? Yes. Viktor could run as one of the agents *behind* Sentrely's control plane. Viktor handles execution; Sentrely enforces what Viktor is allowed to do, logs every action, and gates risky operations on your approval. This is the pattern for teams that want Viktor's ease of use with Sentrely's governance guarantees. ## The Bottom Line Viktor is an excellent product for teams that want an AI coworker with zero setup. But "AI coworker with access to everything" becomes a liability at scale. When you need to answer "what did our agents do, to what, and who approved it?" — that's when you need a control plane. Viktor does the work. Sentrely makes sure the work stays safe. --- ## Blog posts ### Building Your Own Claude Code Gateway: What It Actually Takes **URL:** https://sentrely.com/blog/sentrely-vs-building-your-own **Published:** 2026-04-26 **Description:** You can build a control plane for Claude Code agents yourself. Here's the honest breakdown of what's involved — the components, the effort, and when DIY makes sense. You understand why Claude Code needs a control plane — per-agent credentials, policy enforcement, audit logging, operational controls. Now the question: build it yourself or use an existing solution? Building your own is a legitimate option for some teams. Here's the honest breakdown. ## What You'd Need to Build A functional Claude Code gateway has seven major components. **1. API Proxy Layer (4-6 weeks)** A reverse proxy between Claude Code and the Anthropic API. Intercepts calls, injects session metadata, enforces rate limits, records token usage, applies content filters. Basic proxy: 2-3 weeks. Reliable proxy with connection pooling, retry logic, and graceful degradation: another 2-3 weeks. **2. Credential Vending Service (3-4 weeks)** Generates temporary, scoped credentials per agent session. For AWS: STS AssumeRole with session tags. For Git: short-lived tokens. The AWS integration is straightforward; Git credential vending is trickier (GitHub Apps vs Bitbucket have different token mechanics). **3. Policy Engine (4-6 weeks)** Evaluates whether an operation is allowed for a given session. You can use an existing framework (Casbin, Open Policy Agent, Cedar) or build your own. The policy language design is the hard part — too simple and you can't express real policies, too complex and nobody can write them. **4. Audit Log System (5-7 weeks)** Records every operation with full context in structured, immutable storage. Basic logging to CloudWatch: 2-3 weeks. Query interface, retention policies, export capabilities, and immutability guarantees: another 3-4 weeks. **5. Session Management (3-4 weeks)** Tracks active sessions, their state, resource usage, and provides control mechanisms (pause, resume, kill). The basic lifecycle is simple. Hard parts: real-time cost tracking, graceful session termination, handling sessions that lose connectivity. **6. Slack Integration (4-6 weeks)** Sends approval requests and notifications. Receives callbacks. Routes to the right channels. Basic notifications: 2-3 weeks. Interactive approval workflows, timeout handling, escalation logic: another 2-3 weeks. **7. Dashboard (4-6 weeks)** Web interface showing active sessions, recent operations, cost metrics, policy violations. Needs real-time updates (WebSocket) and historical queries. ## Total Effort | Component | Build | Monthly Maintenance | |---|---|---| | API Proxy | 4-6 weeks | 2-4 hrs | | Credential Vending | 3-4 weeks | 4-6 hrs | | Policy Engine | 4-6 weeks | 4-8 hrs | | Audit Logs | 5-7 weeks | 2-4 hrs | | Session Management | 3-4 weeks | 2-4 hrs | | Slack Integration | 4-6 weeks | 1-2 hrs | | Dashboard | 4-6 weeks | 4-8 hrs | | **Total** | **27-39 weeks** | **19-36 hrs/mo** | That's 6-9 months of a senior engineer's time, plus 20-35 hours per month ongoing maintenance. These are conservative estimates for an engineer who already knows the AWS SDK, has built proxy services before, and understands the Claude Code architecture. ## What You Can Skip Initially Not everything is needed from day one: - **Weeks 1-4:** Proxy + Audit. Get visibility. See what agents are doing, track costs. - **Weeks 5-10:** Policy Engine + Session Management. Add controls. - **Weeks 11-16:** Credential Vending + Slack. Add security and communication. - **Weeks 17-24:** Dashboard. Add visibility for the whole team. This gets you basic visibility in a month and functional governance in four months. ## When DIY Makes Sense Building your own is right when: - **Unusual integration requirements.** Your CI/CD, cloud provider, or internal tools aren't standard. - **Deep customization needed.** Your policy model or approval workflow is unique and can't be configured in existing products. - **Engineering capacity available.** A senior engineer with 6-9 months and no other priorities who will also maintain it ongoing. - **Full code ownership required.** For regulatory or security reasons, you need to own and audit every line. ## When to Use Sentrely Using an existing solution is right when: - **You need governance now, not in six months.** The business value is in the agents, not the infrastructure. - **Standard integrations.** AWS, GitHub/Bitbucket, Slack, standard Claude Code. - **Engineering time is better spent on product.** Every hour building a proxy is an hour not building what customers pay for. - **Compliance deadline looming.** SOC 2 audit in three months? You don't have time to build and prove a custom solution. Sentrely provides all seven components out of the box. The trade-off is the standard build-vs-buy equation: building gives you control and customization, buying gives you speed and reduced maintenance burden. The cost of building isn't just the initial 6-9 months — it's the ongoing 20-35 hours per month of maintenance, the opportunity cost of the engineer's time, and the risk that the engineer who built it leaves and nobody else understands it. ## The Decision Framework Ask three questions: 1. **How urgently do you need governance?** "We needed it yesterday" makes DIY impractical. 2. **How custom are your requirements?** Standard stack → existing solution. Unique requirements → might need to build. 3. **What's the engineering opportunity cost?** Is 6-9 months of a senior engineer on a gateway worth more than whatever else they'd build? For most teams: use an existing solution, ship agents now, build custom components later if you actually need them. The worst option is running Claude Code in production with no governance at all. That's not a build-vs-buy decision. That's a risk acceptance most organizations can't justify. --- ### Claude Code vs GitHub Copilot: Why They Need Different Governance **URL:** https://sentrely.com/blog/claude-code-vs-github-copilot **Published:** 2026-04-26 **Description:** Copilot assists developers while they code. Claude Code agents code without them watching. The difference changes everything about how you govern each. GitHub Copilot and Claude Code both write code. That's where the similarity ends. The difference isn't which model is better or which writes cleaner functions. It's the operating model — and that changes everything about governance, security, and risk. ## The Fundamental Difference **GitHub Copilot is an assistant.** It suggests code completions as you type. You accept, modify, or reject every suggestion. The developer is in the loop on every single line. Copilot never executes anything. It never runs tests, pushes to Git, or creates AWS resources. It suggests text. A human decides what to do with it. **Claude Code is an agent.** You give it a task and it executes. It reads files, writes code, runs shell commands, executes tests, creates commits, interacts with APIs. The developer launches it and steps back. The agent operates autonomously until the task is complete. This distinction — assistant versus agent — is everything for governance. ## Why Copilot Doesn't Need a Control Plane Copilot's security model is the developer. Every suggestion is reviewed before it enters the codebase. The blast radius of a Copilot mistake is one suggestion that a developer accepted. The developer's own permissions limit what that suggestion can do. Normal code review catches what the developer missed. You might add Copilot-specific policies (disable it for certain repositories, configure which models are used). But you don't need a new governance layer. Your existing development governance covers it. ## Why Claude Code Does Need a Control Plane Claude Code operates outside the normal development workflow. It executes commands directly. It can modify files without a pull request. It can run scripts affecting your infrastructure. It creates Git commits and pushes them — all without a human reviewing each action. The blast radius of a Claude Code mistake is whatever the agent has access to. If it has AWS credentials, it can create or destroy resources. If it has Git push access to main, it can deploy code without review. Your existing development governance doesn't cover this because Claude Code doesn't go through your existing workflow. It bypasses code review. It bypasses change management. It bypasses access controls by inheriting the developer's full credential set. ## The Governance Matrix | Concern | Copilot | Claude Code | |---|---|---| | Code review | Normal PR process | Agent bypasses PRs unless forced to branch | | Access control | Developer's IDE permissions | Developer's full credential set | | Blast radius | One code suggestion | Everything the agent can reach | | Audit trail | Git blame shows developer | Need session-level attribution | | Cost control | Flat subscription | Per-token, highly variable | | Kill switch | Close the IDE | Need per-session stop mechanism | | Approval gates | PR review is the gate | Need explicit gates for sensitive ops | Every cell in the Claude Code column represents a governance requirement that doesn't exist for Copilot. "We already govern our AI coding tools" doesn't cover Claude Code. You govern your AI assistant. You haven't governed your AI agent. ## When to Use Each **Copilot:** Interactive coding assistance. You're writing a function, Copilot suggests the implementation. You're in the flow, the human stays in control throughout. **Claude Code:** Autonomous task execution. Implement this feature across three files. Refactor this module. Write tests for this service. The human defines the task and reviews the result — execution is autonomous. Most teams will use both. Copilot for interactive development, Claude Code for batch work and automation. The mistake is treating them the same way. Copilot needs a license and maybe a usage policy. Claude Code needs a control plane. ## The Rule That Covers Everything The more autonomy the tool has, the more governance it needs. A Copilot that only suggests completions needs minimal governance. A Copilot that executes multi-file changes autonomously needs the same governance as Claude Code. The tool name doesn't matter. The operating model does. If your AI coding tool executes commands, modifies files, and interacts with infrastructure without human approval on each action, it's an agent. And agents need a control plane. --- ### Running Claude Code Agents Across a Team: What Changes **URL:** https://sentrely.com/blog/claude-code-teams **Published:** 2026-04-26 **Description:** One developer running Claude Code is manageable. Eight developers running it against shared infrastructure is a completely different problem. Running Claude Code solo is straightforward. You know what the agent is doing because you launched it. You know what it costs because it's your API key. Scale to a team of eight and every assumption falls apart. The shift from single-developer to team usage isn't linear — it's a phase change. Problems that didn't exist at all with one person become the dominant operational concern with eight. ## Credential Management Becomes a Real Problem With one developer, credential management is simple: your API key, your AWS profile, your Git key. The agent uses your identity. The audit trail points back to you. With a team, you have two bad options. Everyone shares one set of agent credentials — no attribution, single blast radius. Or everyone uses their own — credential sprawl, inconsistent access levels. What you actually need is credential vending: a system that generates temporary, scoped credentials for each agent session. Credentials identify the developer, project, and permissions needed. When the session ends, credentials expire automatically. This is standard practice for human access in mature organizations. For Claude Code agents, almost nobody does it yet — that gap is where the security incidents live. ## Blast Radius Multiplies One developer's Claude Code agent can only break what that developer has access to. Eight developers' agents can collectively break everything. Consider: Developer A's agent modifies a shared configuration file. Developer B's agent reads the now-modified config and makes decisions based on it. Developer C's agent detects both changes and tries to "fix" them. Three agents working at cross-purposes against shared infrastructure. This isn't hypothetical — it's what happens when multiple autonomous agents operate in the same environment without coordination. Humans handle this with communication. Agents don't have that social protocol unless you build it in. **Solution:** Workspace isolation and resource locking. Each agent session operates in a bounded scope. If two agents need to modify the same resource, there needs to be a coordination mechanism — a lock, a queue, or a policy preventing concurrent modification. ## Attribution Gets Murky When something breaks, the first question is "who did this?" With team agents sharing credentials: "we don't know." CloudTrail shows the API call came from `staging-deploy-role`. Git log shows the commit was from `ci-bot`. But which developer's agent session made the change? Without session-level attribution you're doing forensics every time something goes wrong. Good attribution requires three things: a unique session identifier that propagates through every system the agent touches, that identifier tied back to the human who launched the session, and the mapping stored in an immutable log. ## Policy Consistency Across Team Members Developer A configures their Claude Code agent with strict guardrails: no production access, branch protection, cost limits. Developer B doesn't bother. Both agents operate against the same infrastructure. Without centralized policy management, your security posture is only as strong as the least careful developer on your team. Organizations use IAM policies instead of trusting humans to self-limit for exactly this reason. The same principle applies to agent access. Centralized policies mean the team doesn't think about security configuration for every session. Policies are defined once, enforced everywhere. A developer can't accidentally run an agent with more permissions than the policy allows. ## Cost Allocation Gets Complicated Eight developers' agent usage is a cost allocation problem. Which team is responsible for the $2,400 Anthropic bill this month? Without per-developer or per-project cost tracking, you can't answer that question. **The fix:** Tag every API call with the developer, project, and session. Anthropic's API supports custom metadata. Your proxy layer should inject it automatically. Your billing dashboard can then slice costs by any dimension you care about. ## The Organizational Shift Running Claude Code across a team isn't a technology problem you solve once. It's an organizational capability you build over time. The team needs to agree on policies, establish credential practices, build attribution into their workflow, and create approval processes that work without creating bottlenecks. This is the same evolution that happened with cloud access, CI/CD pipelines, and infrastructure as code. Each started as a single-developer tool that became a team capability through investment in governance infrastructure. Claude Code is following the same path. The teams that invest in governance early scale smoothly. The teams that don't hit a wall around 4-6 developers where coordination overhead outweighs productivity gains. --- ### Claude Code Security: Permissions, Denylists, and What They Don't Cover **URL:** https://sentrely.com/blog/claude-code-security **Published:** 2026-04-26 **Description:** Claude Code has a real permission system — allow/ask/deny rules, bypass-mode controls, env scrubbing. Here's how it actually works, where the denylist leaks, and why production teams still need an external control plane. Claude Code ships with a permission system. You've probably already disabled it. The `--dangerously-skip-permissions` flag exists because the built-in prompts interrupt agent flow to the point where autonomous operation becomes impractical. So teams skip them — and then they have an agent with no controls running against their infrastructure. That flag name isn't an accident. Anthropic is telling you this is dangerous. The honest version of this article: Claude Code's native permission controls are real and worth configuring — but they're a **partial** control, not a production governance layer. This piece covers exactly how the native system works, the three places it leaks, and where you have to put a control plane instead. ## How Claude Code's permission system actually works Claude Code evaluates every tool call against three rule lists in your `settings.json`: ```json { "permissions": { "allow": ["Bash(npm run test:*)", "Read(src/**)"], "ask": ["Bash(git push:*)"], "deny": ["Read(./.env)", "Read(~/.aws/**)", "Bash(sudo *)", "Bash(rm -rf *)"] } } ``` Two things matter here: 1. **Evaluation order is deny → ask → allow.** A `deny` rule always wins. An `ask` rule prompts a human. An `allow` rule auto-approves. Anything unmatched falls back to the default (interactive prompt). 2. **Settings are layered.** Enterprise *managed* settings override the command line, which overrides project `.claude/settings.local.json`, then `.claude/settings.json`, then user settings. Managed settings can't be overridden by a developer — which is the only reason any of this is enforceable on a team. This is genuinely useful. Deny `Read(./.env)`, `Read(~/.aws/**)`, `Bash(sudo *)`, and `Bash(rm -rf *)` in a managed policy and you've blocked the most obvious foot-guns. **Configure it. It's free and it helps.** But it stops well short of governance, for three reasons. ## Leak #1: `--dangerously-skip-permissions` turns it all off When you skip permissions, every tool call is auto-approved. Every shell command executes. Every file write goes through. The agent has exactly the same access as the user or service account that launched it. If that user has AWS admin credentials, the agent has AWS admin credentials. If their SSH key can push to production, the agent can push to production. If there's a database connection string in the environment, the agent can query or drop tables. You *can* prevent developers from entering bypass mode — set `disableBypassPermissionsMode` to `"disable"` in a managed settings file. But that only works if you ship and enforce managed settings on every machine, and it pushes developers back into interactive prompts, which is the friction they were escaping in the first place. The native system makes you choose between *enforced-but-annoying* and *autonomous-but-ungoverned*. Production needs both. ## Leak #2: the denylist is bypassable by design This is the one most teams miss. A denylist blocks *named paths to a capability*, not the capability itself. Deny `Read(./.env)` and the agent simply runs: ```bash cat .env # Bash, not Read grep -r SECRET . # also reads the file xxd .env | head # so does this ``` Unless you *also* deny `Bash(cat:*)`, `Bash(grep:*)`, `Bash(xxd:*)`… and every other tool that can read a file — which you can't enumerate — the secret is reachable. Anthropic's own permission model evaluates the *declared* tool (`Read` vs `Bash`), so a `Read` deny does nothing about a `Bash` that reads. Security researchers have demonstrated exactly this class of deny-rule bypass. Denylists fail open. The only model that fails *closed* is **deny-by-default**: nothing is allowed unless explicitly permitted, scoped to the specific action and resource. Claude Code's native config is an allowlist-with-holes, not deny-by-default — and you enforce deny-by-default at a layer the agent can't talk around, not in the agent's own config. > Related: [Why static RBAC isn't enough for AI agents](/blog/rbac-for-ai-agents) — agents take dynamic paths to the same outcome, so role lists over-grant or break. ## Leak #3: MCP servers and env vars expand the surface underneath the rules Model Context Protocol servers extend Claude Code's capabilities — a GitHub MCP server creates PRs, a Postgres MCP server runs queries, an AWS MCP server manages cloud resources. There's no permission layer *between* Claude Code and the MCP server: if the server can do it, the agent can do it. Consider a Postgres MCP server you added so the agent can check schema definitions. The server doesn't distinguish `SELECT information_schema` from `DROP TABLE` on production — it's a database connection, and everything the connection can do, the agent can do. "Use a read-only DB user" is correct and insufficient; you need that same least-privilege thinking on *every* MCP server, environment variable, and tool. Environment variables are the quiet version of this. Claude Code can scrub variables from subprocess environments, but by default the agent inherits the shell it launched in — every `AWS_SECRET_ACCESS_KEY`, `DATABASE_URL`, and API token sitting in your env. Scrubbing helps; it's still a per-machine setting you have to ship and verify, not a guarantee. ## The AWS credential problem (a worked example) The most dangerous Claude Code configuration is also the most common: AWS admin credentials, because the developer who set it up uses that profile for daily work. Real scenario: the agent deploys a Lambda. It has `AdministratorAccess`. Mid-deploy it decides the function needs an SQS queue and a DynamoDB table — so it creates them. It decides the function needs internet access — so it modifies a VPC security group. Every action lands in CloudTrail; nobody reviews CloudTrail in real time. Two weeks later the security team finds an over-permissive security group, an unencrypted table, and a queue with no DLQ, and tracing it back to one agent session means correlating timestamps across services. **What proper AWS access looks like** (and what AWS itself now prescribes for AI agents — least-privilege, actor-type-differentiated IAM): - Temporary credentials from STS `AssumeRole`, not long-lived keys - IAM policies scoped to specific resources (this bucket, this function) — `s3:GetObject`, not `s3:*` - Service-level restrictions (read EC2, cannot modify security groups) - Session duration limits (expire in 1 hour, not 12) - Session tags identifying the agent, developer, and project That's a per-request decision, not a static profile. See [Governing Claude Code's AWS access](/blog/claude-code-aws) for the full pattern. ## When "no controls" actually bites: the `terraform destroy` story Early 2026, a team running Claude Code in auto mode let it "clean up unused infrastructure." It ran `terraform destroy` against a workspace pointed at production and wiped roughly two and a half years of data before anyone noticed. No deny rule for `Bash(terraform destroy*)` existed because nobody thought to write it — which is the whole problem with denylists: you only block the disasters you predicted. A deny-by-default policy with a **human approval gate** on destructive infra commands (`terraform apply/destroy`, `kubectl delete`, `rm -rf`, `sudo`, `aws … delete-*`) turns that incident into a one-tap "deny" in Slack. See [approval gates for Claude Code](/blog/claude-code-approval-gates). ## What a real security layer looks like The model production needs has four components, and only the first is partially covered by native config: - **Identity** — every session attributed to a specific human and purpose; session ID in every log line; any action traceable to who launched it and why. - **Policy** — every operation checked against a deny-by-default engine *before* it executes, defined by the team (not the individual developer), scoped to action + resource, at a layer the agent can't reconfigure. - **Audit** — every operation in an immutable, structured log: which session, which developer, what operation, what parameters, what result. Not "agent called AWS API." - **Control** — stop any session instantly, change policy without restarting agents, set cost limits, and require human approval for specific operations. Native Claude Code config gives you a weak version of policy and nothing durable for identity, audit, or control. The gap between "Claude Code with no controls" and "Claude Code with production security" is exactly where a [control plane](/blog/what-is-an-ai-agent-control-plane) sits: a gateway between the agent and your infrastructure that enforces policy, logs every action immutably, and gives you a kill switch. In that model the agent holds **zero credentials** — it asks the gateway, the gateway decides, the gateway acts. ## The uncomfortable truth Claude Code's security model assumes a human is watching. `--dangerously-skip-permissions` is an escape hatch for development, not a production configuration, and the native permission system — useful as it is — fails open: skippable, bypassable, and per-machine. Without an external enforcement layer, the answer to every question about agent permissions is "the agent can do whatever its credentials allow." That's not security. That's hope. ## FAQ **Does Claude Code have built-in security?** Yes — a permission system with allow/ask/deny rules (deny takes precedence), layered settings, a `disableBypassPermissionsMode` managed control, and subprocess env scrubbing. It's worth configuring, but it's a partial control: it can be skipped with `--dangerously-skip-permissions`, its denylists are bypassable (a `Read(./.env)` deny doesn't stop `Bash(cat .env)`), and it's enforced per-machine rather than centrally. **Is `--dangerously-skip-permissions` safe for production?** No. It auto-approves every tool call and gives the agent the full access of the account that launched it. Use it only in disposable sandboxes; in production, route the agent through a policy-enforced gateway instead. **How do I stop Claude Code from reading secrets or running destructive commands?** Native deny rules (`Read(~/.aws/**)`, `Bash(sudo *)`, `Bash(rm -rf *)`) help but are incomplete because the same capability has many command paths. A deny-by-default control plane with human approval gates on destructive operations is the enforceable answer. **What's the difference between Claude Code permissions and an AI agent control plane?** Permissions live in the agent's own config (skippable, editable by the developer, per-machine). A control plane sits *outside* the agent as a proxy it can't reconfigure — enforcing centralized policy, holding the credentials so the agent has none, and producing an immutable audit trail. --- ### Claude Code in Production: The 5 Things That Break First **URL:** https://sentrely.com/blog/claude-code-in-production **Published:** 2026-04-26 **Description:** Moving Claude Code from local development to production automation is harder than it looks. Here's exactly what breaks and how to prevent each failure. Claude Code running on your laptop feels magical. You describe a task, it executes, the work gets done. Then you move it to production — against real infrastructure, with real credentials, running autonomously — and things break in ways nobody anticipated. The failure modes are never dramatic. No explosion. Just slow-building problems that become expensive before anyone notices. After watching dozens of teams attempt this transition, here are the five things that consistently break first. ## 1. Shared Credentials Across the Team Every developer uses the same Anthropic API key. Works fine in development. Then you move to production, and twelve Claude Code agent sessions are running against your AWS account — all authenticated with the same IAM user. When something goes wrong, you open CloudTrail and every action reads "performed by claude-agent-prod." Which agent? Which developer triggered it? Which session? Useless. The real scenario: A Claude Code agent creates an overly permissive security group rule. Your security team flags it two days later. You spend half a day tracing it because every agent operation uses the same credential identity. **Prevention:** Every agent session needs its own credential identity. Use AWS STS AssumeRole with session tags that identify the developer, project, and session ID. Your audit logs should trace any action back to the specific human who triggered it. ## 2. No Branch Protection for Agents Your developers need pull requests to merge to main — but nothing enforces that rule for Claude Code. The agent has service account Git credentials. Branch protection rules often don't apply to service accounts. The agent pushes directly. The real scenario: A developer sets up a Claude Code agent for automated dependency updates. It encounters a breaking change, writes a fix with a subtle bug, and pushes directly to main. The broken code deploys automatically. **Prevention:** Agents should use Git credentials subject to the same branch protection as humans. Configure policies that force agents to work on feature branches and create pull requests. The PR is the review gate — don't let agents bypass it. ## 3. Context Window Exhaustion on Large Codebases Claude Code works by ingesting your codebase into its context window. On a production codebase with hundreds of files, the agent silently degrades. It doesn't crash. It just starts making worse decisions — missing the utility function that already exists, writing duplicates, ignoring naming conventions from other modules. Nobody realizes the agent is context-limited until code review catches three duplicate implementations. **Prevention:** Monitor context utilization per session. Break large tasks into smaller, focused sessions. Point the agent at specific files rather than letting it ingest everything. ## 4. No Cost Visibility Claude Code uses tokens. Every file read, every tool call, every retry loop — tokens. Without visibility, a single stuck session can consume more tokens than a week of normal usage before anyone notices. The real scenario: An agent gets stuck in a retry loop on Friday afternoon. By Monday morning it has consumed $400 in API credits. Nobody knew because there was no real-time cost tracking. **Prevention:** Implement per-session token budgets with hard stops. Track usage in real time. Alert at 50% and 80% of budget. When the budget is hit, the session stops — no exceptions. ## 5. Missing the Kill Switch When a Claude Code agent goes off the rails, how fast can you stop it? "SSH in and kill the process" is too slow. "Revoke the API key" kills every agent, not just the problematic one. The real scenario: An agent with S3 write permissions starts "organizing" a bucket — moving objects based on its interpretation of project structure. The interpretation is wrong. Objects move to locations where other services can't find them. Finding and stopping the right process takes 15 minutes. The agent moves 200 more objects in that time. **Prevention:** Every agent session needs a unique session ID and a kill endpoint. Build a Slack command or dashboard that lists running sessions and lets you stop any one instantly. Test the kill switch before you need it. --- The pattern: all five problems are infrastructure gaps, not AI capability gaps. Claude Code itself works fine. What breaks is everything around it. The teams that succeed in production invested in the control plane — the layer between the AI agent and your infrastructure that enforces policies, tracks costs, maintains audit trails, and provides operational control. --- ### Claude Code Cost Control: Why Your Bill Is Higher Than It Should Be **URL:** https://sentrely.com/blog/claude-code-cost-control **Published:** 2026-04-26 **Description:** Uncontrolled Claude Code agents are expensive. Here's exactly where your token budget goes and how to cut it 40-60% without losing any capability. Claude Code is expensive. Not inherently — the per-token pricing is reasonable. It's expensive because of how tokens get consumed in practice, and most teams have no visibility into where the money goes until the invoice arrives. A typical Claude Code session costs $0.50–$5.00 for focused, well-scoped work. But sessions routinely hit $20, $50, even $100+ when they go off track. With agents running across a team, those numbers multiply fast. ## Where Your Tokens Actually Go **Context accumulation.** Every message includes the full conversation history. The first tool call might cost 500 tokens. The fiftieth costs 500 tokens for new content plus 25,000 tokens of accumulated context. By the end of a long session, you're paying mostly for history, not new work. A session with 100 tool calls isn't 100x the cost of one call — it's closer to 500x because of context growth. **System prompt overhead.** Claude Code's system prompt includes project structure, CLAUDE.md instructions, and tool definitions — easily 3,000–5,000 tokens. Sent with every single API call. Over 80 calls, that's 240,000–400,000 tokens just for the system prompt. **Retry loops.** The agent makes a change, runs the test, the test fails, it analyzes, makes another change, runs the test again. Each cycle is a full API roundtrip with accumulated context. Three retry cycles cost more than the initial implementation. **File reading.** When Claude Code reads a file, the contents go into context. Read ten files and you've added thousands of tokens to every subsequent call in the session. Real numbers from one team's monthly breakdown: 42% of spend went to context re-transmission, 19% to tool output (file reads, test results), 16% to system prompt overhead, 13% to retry loops. Only 10% went to the agent actually reasoning about the problem. ## How to Cut 40-60% **Break work into smaller sessions.** Instead of one 100-tool-call session implementing a feature end-to-end, use three 25-tool-call sessions: plan, implement, test. Each starts with fresh context. You lose some continuity but eliminate the context accumulation that compounds costs exponentially. A 100-tool-call session costs roughly 3-4x what the same work costs split across four 25-tool-call sessions. **Set session budgets with hard stops.** Configure a maximum token budget. When reached, the session stops. This is critical — a soft limit that sends a warning but continues is useless. The agent doesn't read warnings. The stop must be hard. **Implement circuit breakers for retry loops.** If the agent has tried the same task three times and the test still fails, stop the session. The problem likely requires a different approach or human input. Letting it try a fourth, fifth, sixth time rarely succeeds and always costs more. **Scope the context intentionally.** Instead of letting Claude Code read your entire codebase, tell it which files are relevant. "Read src/auth/middleware.ts and src/auth/types.ts, then implement the new auth check" is cheaper than "look at the codebase and figure out how auth works." **Trim tool output.** Configure verbose test runners for minimal output inside Claude Code sessions. The agent needs to know which test failed and why — not the full 200-line diff and stack trace. ## Tracking Costs at the Right Granularity If you only see total monthly spend, you can't identify which sessions are expensive or why. Track at three levels: **Per-session:** How much did this specific run cost? Catches runaway sessions immediately — a session that cost $40 when similar sessions cost $2 is immediately visible. **Per-developer:** Which team members are consuming most tokens? Identifies training opportunities — maybe one developer's agent sessions are structured inefficiently. **Per-project:** Which projects cost most? A monorepo project that costs 10x more than others probably needs to be split so agent sessions work with smaller codebases. ## After Optimization The same team that had only 10% productive reasoning implemented session splitting, context scoping, and circuit breakers. Their next month: - Total cost dropped from $2,000 to $880 (56% reduction) - The amount spent on actual productive reasoning increased slightly (from $200 to $210) - Retry loop cost dropped from $260 to $40 Shorter sessions means less context noise for the agent to process. The agent actually performs better while costing dramatically less. ## The Bigger Lever: Stop Renting Tokens Optimization trims a metered bill — but the question worth asking first is who has $10k/mo to spend on agent tokens at all. If your team runs Claude Code or Codex agents all day, a per-token API meter climbs into the thousands fast. The structural fix is to stop renting tokens by the hour. Connect your agents via OAuth to the **Claude Max** or **ChatGPT / Codex** subscription you already pay for — a flat ~$100–200/mo — instead of a metered API key, and you get the same horsepower without the runaway invoice. Sentrely supports both: an API key when you want raw per-token control, or BYO subscription via OAuth when you want a predictable flat rate. Bring your plan, not your token bill. The worst thing you can do is run Claude Code without cost visibility. Without it, you can't tell the difference between a $2 session and a $200 session until the bill arrives. --- ### Claude Code and SOC 2: What Auditors Actually Ask **URL:** https://sentrely.com/blog/claude-code-compliance **Published:** 2026-04-26 **Description:** Your SOC 2 auditor will ask about your AI agent controls. Here's exactly what they want to see and how to show it — based on real audit conversations. Your SOC 2 auditor is going to ask about Claude Code. Not if. When. If you're using it in your development process — especially if it touches production systems — the auditor needs to understand your controls. SOC 2 doesn't have a specific "AI agent" section. But agents touch multiple Trust Service Criteria: Security, Availability, Processing Integrity, and Confidentiality. The auditor maps your agent usage to existing criteria and expects controls that satisfy each one. Here's what they actually ask and what they want to see. ## "Describe Your AI Agent Access Controls" Usually the first question. The auditor wants to understand who can deploy agents, what agents can access, and how access is controlled. **What "we use Claude Code" sounds like:** "We have an autonomous system with access to our production infrastructure. We don't know exactly what it can do. Different developers configure it differently." **What they want to hear:** "Each Claude Code agent session launches with a specific IAM role scoped to the task. Sessions are attributed to the launching developer. We use temporary credentials that expire after one hour. The policy engine prevents operations outside defined scope." **Artifacts they'll request:** access control matrix showing which roles/agents can access which systems, policy definitions (the actual configuration), evidence of policy enforcement (denied operations, not just allowed ones). ## "How Do You Log AI Agent Activity?" SOC 2 Control CC7.2 requires monitoring system components for anomalies. An AI agent modifying your infrastructure is exactly the kind of component that needs monitoring. **What they expect to see:** Structured logs capturing every agent operation with timestamps, session identifiers, the developer who initiated the session, the operation performed, the target resource, and the result. They'll pick random entries and ask you to trace them back to a specific developer and task. **What fails:** API usage logs from Anthropic (shows token consumption, not actions), application logs without session correlation, logs the dev team can modify or delete, missing time periods where agents were active but no logs exist. **What passes:** Immutable, structured audit logs in write-once storage; every entry tied to a session ID, developer, and project; retention matching your compliance requirement (usually 1 year minimum); ability to produce a complete timeline for any session within minutes. ## "What Is Your Change Management Process for Agent-Initiated Changes?" SOC 2 Control CC8.1 covers change management. When Claude Code modifies production code or infrastructure, that's a change. The auditor needs to see that changes go through your defined process regardless of whether a human or agent initiated them. This is where most teams fail. Their change management process assumes human actors: developer writes code, submits PR, reviewer approves, CI/CD deploys. Claude Code can bypass all of these if it has direct push access and production credentials. **What they want:** Agent-initiated changes go through the same approval process as human-initiated changes. Code changes submitted as pull requests, not pushed directly. Infrastructure changes require authorized approval. Emergency changes have a defined exception process. **Evidence they'll request:** Pull requests created by Claude Code agents (showing review and approval), approval records for infrastructure operations, denied change requests (showing controls work in both directions). ## "How Do You Handle Incidents Involving AI Agents?" SOC 2 Control CC7.3 requires incident response procedures. The auditor will ask: if a Claude Code agent causes an incident, how do you detect, respond, and prevent recurrence? **Your incident response plan should cover:** *Detection:* Real-time monitoring of agent operations, automated alerts for anomalous behavior. *Containment:* Per-session kill mechanisms, credential revocation, ability to halt all agent operations. *Investigation:* Session-level audit trails with ability to replay the agent's decision-making process. *Prevention:* Policy updates, scope restrictions, additional approval gates for the operation type that caused the incident. **The auditor will ask for evidence of at least one incident or near-miss handled through this process.** If you've had no incidents, they'll ask about testing: do you run tabletop exercises? Test your kill switch? Verify policy denials work? ## "How Do You Ensure Data Confidentiality with AI Agents?" If Claude Code processes sensitive data, the auditor scrutinizes data handling controls under CC6.1 and CC6.5. This is a real concern: Claude Code reads files and sends their content to Anthropic's API for processing. If those files contain customer PII, financial data, or health records, that data leaves your environment. **Controls the auditor wants:** - Data classification identifying sensitive files/directories - Agent policies preventing access to classified data - Network controls restricting where agent sessions can send data - For highly sensitive environments: VPC-deployed models keeping all data internal ## "Show Me Your Vendor Risk Assessment for Anthropic" Anthropic is a vendor. SOC 2 requires vendor risk assessment under CC9.2. Many teams have thorough assessments for their database provider and payment processor, and none for their AI provider. **They'll ask:** Have you reviewed Anthropic's SOC 2 report? Do you have a data processing agreement? What happens to your data after API processing? What's your contingency plan if the API is unavailable? ## Preparing for the Audit Checklist if your SOC 2 audit is coming: 1. **Document your agent access control model.** Write down which roles launch agents, what each can access, how policies are enforced. Include the actual policy files. 2. **Verify your audit trail.** Pull logs for 90 days. Can you trace any random agent session from launch to completion? 3. **Review your change management.** Are agent-initiated code changes going through pull requests? Infrastructure changes through approval gates? 4. **Update your incident response plan.** Add agent-specific sections. Run at least one tabletop exercise. 5. **Complete a vendor risk assessment for Anthropic.** Review their SOC 2, document your data handling controls. 6. **Test your controls.** Try to violate a policy — does the agent get denied? Try to kill a session — does it stop immediately? The teams that pass SOC 2 with Claude Code aren't doing anything exotic. They're applying the same governance principles they use for human access — identity, authorization, logging, monitoring, incident response — to their agent operations. The principles are the same. The tools are different. What they're not doing is running Claude Code with admin credentials and no logging, hoping the auditor doesn't ask. The auditor will ask. --- ### Secure AI Agent Access to AWS: Least Privilege for Claude Code **URL:** https://sentrely.com/blog/claude-code-aws **Published:** 2026-04-26 **Description:** How to govern an AI agent's AWS access the way AWS itself recommends — least-privilege IAM, temporary credentials, actor-differentiated policies, and a control plane so Claude Code holds zero long-lived keys. Claude Code is remarkably effective at AWS operations — it writes CloudFormation, debugs Lambdas, manages S3, configures ECS, and troubleshoots networking. Give it credentials and a task and it figures out the API calls. That's the appeal and the danger. "Figure out the API calls" means the agent uses whatever permissions you gave it to accomplish the task most efficiently. If the shortest path is a public S3 bucket, an open security group, or an expensive instance type — the agent does it. Not maliciously. Efficiently. AWS now has a published position on this. Its 2026 Security Blog on AI-agent access patterns is blunt: **"You must assume an agent can do anything within its granted entitlements."** The guidance is least-privilege, short-lived credentials, and — critically — **differentiating AI-driven from human-initiated actions** with different IAM rules per actor type. This article is how to implement that for Claude Code. ## The admin-credential anti-pattern The most common Claude Code + AWS setup: the developer exports their `AdministratorAccess` profile into the agent's environment. Every API is available; every resource is reachable. What follows, repeatedly: - **The cleanup incident.** An agent deploying a new service notices "unused" resources and deletes an old Lambda (processing webhooks), an "empty" S3 bucket (nightly DB backups), and an EC2 instance (a bastion host). - **The cost escalation.** An agent "optimizing" RDS resizes it to `db.r6g.4xlarge`. Monthly cost goes $400 → $3,200. Nobody notices for two weeks. - **The security exposure.** An agent debugging connectivity adds an inbound rule for `0.0.0.0/0` on 443. Problem solved; public hole created. None are Claude Code bugs. The agent did exactly what it thought was needed. Nothing constrained its scope. ## Least privilege for Claude Code (AWS's own prescription) **Start with the task, not the service.** An agent deploying a Lambda needs `lambda:UpdateFunctionCode` on that specific function ARN — not `lambda:*` on `*`. AWS's example is identical: grant `s3:GetObject`, not `s3:*`. **Scope to specific resources:** ```json { "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"], "Resource": "arn:aws:s3:::my-app-deployments/builds/*" } ``` This agent reads/writes one prefix of one bucket. It cannot list the bucket, delete objects, change bucket policy, or touch any other bucket. **Separate read from write.** Most agent tasks (log analysis, config review, status checks) need read only. Keep separate read-only and read-write roles; default to read-only; escalate only when the task requires it. **Constrain with conditions:** ```json { "Condition": { "StringEquals": { "ec2:InstanceType": ["t3.micro", "t3.small", "t3.medium"] } } } ``` The agent can launch EC2 — but only small instances. A `p4d.24xlarge` just returns a permission error, invisible to the agent. **Differentiate the actor.** This is the part most teams skip and AWS specifically calls out: an AI-driven action should not run under the same IAM permissions as the human who launched the session. Tag the session as agent-initiated and apply tighter rules to it — so "Jordan running a command" and "Jordan's agent running a command" are governed differently. ## Temporary credential vending (no long-lived keys) Long-lived access keys are the biggest IAM anti-pattern for Claude Code. Use STS `AssumeRole` session credentials instead: when a session starts, a vending service mints temporary creds with a short duration (1 hour is usually enough); when the session ends or they expire, access stops automatically. **Session tags make attribution automatic:** ```json { "Tags": [ {"Key": "agent-session", "Value": "sess_a8f3b21c"}, {"Key": "developer", "Value": "jordan"}, {"Key": "project", "Value": "backend-api"} ]} ``` Now every CloudTrail entry reads `assumed-role/agent-deploy-role/sess_a8f3b21c` with tags identifying the session, developer, and project. ## When things go wrong anyway Even with scoped IAM, agents do unexpected things *within* their allowed permissions. - **Real-time CloudTrail monitoring.** Fire event rules on agent-tagged sessions performing sensitive actions — IAM changes, security-group edits, deletions. - **Session kill switch.** `aws sts revoke-session` (revoke-session-credentials) instantly invalidates a role session's temporary credentials; the agent's next call fails. Practice this before you need it. - **Automated rollback.** AWS Config rules detect drift (e.g., an opened security group) and auto-remediate. ## The gateway approach: zero credentials in the agent All of the above — role management, credential vending, CloudTrail alerting, kill switches — can be done by hand. It stops scaling the moment you have multiple developers running multiple agents across multiple AWS accounts. A [control plane](/blog/what-is-an-ai-agent-control-plane) centralizes it. The agent never receives AWS credentials. It makes requests through a gateway that: 1. Checks the request against a deny-by-default policy engine 2. Injects temporary, scoped credentials for that one operation 3. Logs it with full session context (immutable) 4. Enforces rate limits, cost controls, and [human approval](/blog/claude-code-approval-gates) for sensitive actions The agent never sees the credentials — it can't leak them, cache them, or use them outside policy. **The IAM policy is the floor; the gateway policy is the ceiling.** This is also the cleanest way to implement AWS's "differentiate the actor" guidance: the gateway *is* the agent's identity, distinct from the human's. ## Getting started Running Claude Code against AWS today without controls? Five standard-AWS steps that cut blast radius immediately: 1. **Stop using admin credentials** — create a task-specific role with only what the agent needs. 2. **Switch to temporary credentials** — STS `AssumeRole`, not access keys. 3. **Add session tags** — identify every API call by session, developer, project. 4. **Set up CloudTrail alerting** — know when agent sessions do sensitive things. 5. **Test the kill switch** — practice `revoke-session` before you need it. Then decide whether to operate all of this yourself or route through a managed control plane that does it on every request. ## FAQ **How do I give an AI agent secure access to AWS?** Use least-privilege IAM scoped to specific actions and resource ARNs, temporary STS credentials (not access keys), session tags for attribution, and a separate agent identity from the human's. AWS's own guidance is to assume the agent can do anything its entitlements allow — so scope tightly and differentiate AI-driven from human actions. **What IAM permissions should a Claude Code agent have?** Only what the specific task needs — e.g. `s3:GetObject` on one prefix, `lambda:UpdateFunctionCode` on one function ARN. Default to read-only; escalate per task; never `*:*`. **How do I stop an AI agent that's misbehaving in AWS?** Revoke its STS session (`aws sts revoke-session-credentials`) to instantly kill all its temporary credentials, or — if it routes through a control plane — flip the gateway kill switch, which stops it without touching IAM. **Should the AI agent use my AWS credentials?** No. Long-lived keys in the agent environment are the top anti-pattern. Vend short-lived scoped credentials per session, or hold the credentials in a gateway so the agent has none at all. --- ### Claude Code Approval Gates: When to Let Agents Run and When to Stop Them **URL:** https://sentrely.com/blog/claude-code-approval-gates **Published:** 2026-04-26 **Description:** Not every Claude Code operation needs human approval. Building the right model is the difference between useful automation and security theater that everyone ignores. The first instinct when deploying Claude Code in production: approve everything. Every file write, every command, every Git operation — a human reviews first. This lasts about one day before the team realizes they've created a system requiring more human effort than doing the work manually. The second instinct: approve nothing. Let the agent run freely. This works until something goes wrong, and then the conversation becomes about why nobody was watching. The right answer is in between, and finding it requires understanding what actually needs oversight. ## The Approval Taxonomy Every Claude Code operation falls into four categories: **Always allow.** Operations with no blast radius and easy reversibility: reading files, running tests in a sandbox, listing directories, checking Git status, querying read-only APIs. Gating these creates friction with zero security benefit. **Allow with logging (notify gate).** Operations with minor side effects within expected scope: writing files in the working directory, creating Git commits on a feature branch, running linters, installing development dependencies. Log and make visible, but don't stop the agent. **Allow with approval (approval gate).** Operations with significant side effects on shared resources: pushing to Git, creating pull requests, modifying infrastructure, writing to production databases, sending external communications. A human confirms before proceeding. **Never allow (deny gate).** Operations exceeding the agent's intended scope: pushing to main/production branches, deleting infrastructure, modifying IAM policies. Blocked by policy regardless of approval. If the agent tries, that's an alert. ## The commands that always need a gate Some operations are dangerous enough that they should never run without a human in the loop, no matter how confident the agent is. Practitioners converge on the same short list — gate (or hard-block) these by default: - `rm -rf`, `chmod`, and filesystem destruction - `sudo` and privilege escalation - `ssh` / `scp` to other hosts - `kubectl apply` / `kubectl delete` - `terraform apply` / `terraform destroy`, `cdk deploy` / `cdk destroy` - `aws … delete-*`, security-group, and IAM changes - package installs (`npm install`, `pip install`) that pull arbitrary code The cost of skipping this is not hypothetical. In early 2026 a team let Claude Code "clean up unused infrastructure" in auto mode; it ran `terraform destroy` against a workspace pointed at production and wiped roughly two and a half years of data. There was no gate on `terraform destroy` — because nobody predicted that exact command. That's the trap with hand-written denylists: you only block the disasters you imagined. An approval gate on the *category* (destructive infra) turns that incident into a one-tap "Deny" in Slack instead of a post-mortem. This is why the gate belongs at a [control plane](/blog/what-is-an-ai-agent-control-plane) the agent can't reconfigure, not in the agent's own settings. ## Implementing Approval Gates in Slack Slack is the natural approval interface because it's where engineers already work. A well-implemented gate: 1. Agent hits an operation requiring approval 2. Gateway sends Slack message with full context: operation, agent session, developer who launched it, why the agent wants to do it 3. Authorized approver clicks "Approve" or "Deny" 4. Gateway receives callback, allows or denies the operation 5. Decision logged with approver identity and timestamp **What the message should include:** ``` Agent Request: Push to feature/auth-refactor Session: sess_a8f3b21c (jordan@company.com) Project: backend-api Files changed: 4 (auth/middleware.ts, auth/types.ts, tests/auth.test.ts, README.md) Commit: "Refactor auth middleware to use session tokens" [Approve] [Deny] [View Diff] ``` The "View Diff" button lets the approver review changes without leaving Slack. The approver has enough context to make an informed decision in under 30 seconds. ## Avoiding Approval Fatigue The biggest risk isn't too few gates — it's too many. When every operation requires approval, approvers click "Approve" without reading. This is worse than no gates: it creates false security while adding latency. **Signs of fatigue:** Average approval time under 10 seconds (nobody is reading), approval rate over 99% (gates aren't catching anything), developers complaining agents are too slow, approvers auto-approving from mobile notifications. **Prevention:** Gate only operations where denial is a realistic outcome. Group related operations into a single approval (approve the entire push, not each file individually). Track metrics on approval times and rates — deteriorating numbers indicate fatigue. ## Timeout Behavior When nobody responds to an approval request, what happens? **Block forever:** Agent waits indefinitely. The work stalls completely if the request comes in after hours. **Auto-approve after timeout:** Undermines the entire point. Attackers just wait. **Auto-deny after timeout (recommended):** The operation is denied, the agent is notified and can attempt an alternative approach or stop. Fail closed. This is the right default. **Escalate then deny:** Primary approver gets 15 minutes. No response → escalate to secondary. No response within 30 minutes → auto-deny. Balances responsiveness with security. ## Building Toward More Autonomy Approval gates should evolve as confidence in the agent increases. **Phase 1:** Gate all write operations, all Git pushes, all infrastructure changes. You're learning what the agent does. **Phase 2:** After a month, review approval logs. Operations with 100% approval rate → convert to notify gates. Operations that were denied → keep the gates, they're working. **Phase 3:** Different policies for different projects. The internal tooling project needs minimal gates. The production API needs strict gates. **Phase 4:** Track the agent's history per project. 50 successful sessions with zero incidents and zero denials → consider loosening gates. An incident → tighten and require additional approval for that operation type. ## A Mature Setup For a team of eight with well-tuned policies: - Read operations: no gate (100+ per session) - Feature branch commits: notify (10-20 per session) - Feature branch pushes: approval (2-3 per session) - Pull request creation: approval (1 per session) - Production infrastructure: approval from team lead (rare) - Direct push to main: denied (never) The agent runs at full speed for 95% of operations. The 5% requiring approval are meaningful decisions where a human check adds real value. Approvers handle 3-5 requests per session — manageable without fatigue. This is the balance. Not security theater where everything is gated and everyone rubber-stamps. Not cowboy mode where nothing is controlled. Thoughtful gates at the right decision points, with a path toward more autonomy as confidence grows. ## FAQ **Which AI agent commands should require human approval?** Gate operations that are high-blast-radius or hard to reverse: production deploys, `terraform apply/destroy`, `kubectl delete`, `rm -rf`, `sudo`, IAM/security-group changes, data deletion, external communications, and financial transactions. Let read-only and sandboxed operations run freely so the gate stays meaningful. **How do human-in-the-loop approvals work for Claude Code?** The agent hits a gated operation; the control plane pauses it and sends an approval request — with full context — to Slack, Telegram, or a dashboard; an authorized human approves or denies; the decision is logged with identity and timestamp. Fail closed: if no one responds, auto-deny. **Won't approval gates slow the agent down?** Only if you over-gate. Gate the ~5% of operations where denial is a realistic outcome; let the other 95% run. A well-tuned setup surfaces 3–5 approvals per session — meaningful checks, not rubber-stamps. **What happens if nobody approves in time?** Auto-deny (fail closed) is the correct default — auto-approving on timeout defeats the purpose, since an attacker just waits. Optionally escalate to a secondary approver before denying. --- ### Building an Audit Trail for Claude Code Agents: What You Actually Need **URL:** https://sentrely.com/blog/claude-agents-audit-trail **Published:** 2026-04-26 **Description:** When your security team asks 'what did your Claude agent do last Tuesday?' — can you answer? Here's what a real agent audit trail looks like versus what most teams have. Your security team is going to ask about your Claude Code agents. Not if — when. If your answer is "we have Anthropic API usage logs," you're going to have a bad meeting. API usage logs tell you tokens were consumed. They don't tell you what the agent did with those tokens. They don't show the shell commands executed, the files modified, the Git commits pushed, or the AWS resources created. The actual actions — the things auditors care about — happen downstream of the API call. ## What Needs to Be in the Log A complete agent audit trail captures five categories: **Session lifecycle.** When the session started, who launched it, what project, what permissions were granted, when it ended, and why (completed, timed out, killed, budget exhausted). **Tool invocations.** Every tool call — shell commands, file reads and writes, MCP server calls, API requests. For each: tool name, input parameters, output or error, timestamp. **Infrastructure operations.** AWS API calls, Git operations, database queries, HTTP requests to external services. These have side effects on external systems — they need special attention because they're the ones that cause damage. **Policy decisions.** Every time an action was checked against a policy — what was requested, what rule applied, whether it was allowed or denied. Denials are especially important: they show your controls are working. **Cost metrics.** Token consumption per tool call, cumulative session cost, budget thresholds crossed. Cost events are a leading indicator — a session that suddenly spikes in cost is often stuck in a loop. ## What Format Satisfies SOC 2 Auditors SOC 2 auditors care about three things: completeness, immutability, and retrievability. **Completeness:** Every relevant event is captured. If an agent modified a production database, that event must be in the log. Gaps where the agent was active but no logs exist will be flagged. **Immutability:** Logs can't be altered after the fact. This means write-once storage. S3 with Object Lock, CloudWatch Logs (append-only), or a dedicated log management service that enforces immutability. **Retrievability:** You can find what you need. If the auditor asks for everything session X did on April 15th, you produce that in minutes, not days. This requires structured logs (JSON, not free-text), consistent session IDs across all entries, and a query mechanism. A practical event format: ```json { "timestamp": "2026-04-15T14:32:07Z", "sessionId": "sess_a8f3b21c", "developerId": "dev_jordan", "projectId": "proj_backend", "eventType": "tool_invocation", "tool": "bash", "input": "aws s3 cp config.json s3://prod-config/app/config.json", "output": "upload: ./config.json to s3://prod-config/app/config.json", "policyResult": "allowed", "policyRule": "s3:PutObject on prod-config/app/*" } ``` Every event ties back to a session, developer, and project. Every event has a policy result. This is what auditors want to see. ## What a Bad Audit Trail Looks Like The most common failure mode is logging at the wrong level of abstraction. **Bad:** ``` 2026-04-15 14:32:00 INFO anthropic.api POST /v1/messages 200 1247 tokens 2026-04-15 14:32:05 INFO anthropic.api POST /v1/messages 200 893 tokens ``` This tells you nothing. What did the agent do between those API calls? What files did it read? What commands did it execute? The Anthropic API logs are the conversation. The tool execution logs are the actions. Auditors care about the actions. Another failure: logging without session correlation. If you have logs from Claude Code, CloudTrail, and your application but they share no common identifier, correlating them during an incident is manual detective work. Every log entry needs a session ID. ## Retention and Export SOC 2 typically requires 1 year minimum. HIPAA requires 6 years. Practical approach: store recent logs (90 days) in a queryable system like CloudWatch Logs or Elasticsearch. Archive older logs to S3 with Object Lock for the compliance retention period. Build an index to find archived logs by session ID, developer, or date range. Your audit system should support three export paths: 1. **Real-time streaming** to your SIEM for security monitoring and alerting 2. **Periodic export** to your compliance archive (daily or weekly) 3. **On-demand export** for incident response — when someone asks for everything related to a production incident, you produce a complete, chronological report filtered by session IDs ## Where to Log The critical decision is where in the stack to log. Logging inside the agent process is unreliable — if the agent crashes, the last few events might be lost. Logging at the gateway layer is more reliable because the logging infrastructure is independent of the agent lifecycle. This is one reason gateway architectures exist. The gateway sees every operation the agent performs and logs it before the operation reaches the infrastructure. If the agent crashes, the logs are intact. If the agent tries something outside its policy, the denial is logged even though the operation didn't execute. Your auditors will thank you. More importantly, the first time something goes wrong — and it will — you'll thank yourself. The bar is also rising specifically for MCP-based agents: as agents reach tools and data through MCP servers, "MCP audit logging" is becoming its own compliance expectation (HIPAA, GDPR, SOX, PCI-DSS), and a generic API log won't satisfy it. The differentiator that holds up under scrutiny is an **immutable, legal-grade** trail with **per-agent identity** — not just "we log things," but "here is the tamper-proof record of exactly which agent did what, six months ago." ## FAQ **What is an AI agent audit trail?** An immutable, chronological record of everything an agent did — every tool call, command, file change, infrastructure operation, policy decision, and approval — each tied to a session, agent identity, and timestamp. It is *not* API usage logs, which only show tokens consumed; the actions downstream of the API call are what auditors care about. **Does Claude Code produce an audit log?** Not a complete one. Anthropic API logs show token usage, not the shell commands, file writes, git pushes, or AWS calls the agent actually made. A real audit trail is captured at the gateway layer, independent of the agent process. **What do SOC 2 and HIPAA auditors require for AI agents?** Completeness (every relevant action logged), immutability (write-once — e.g. S3 Object Lock), and retrievability (structured, session-correlated, queryable). Retention: SOC 2 ≥ 1 year, HIPAA ≥ 6 years. The same expectations now extend to MCP-based agent actions. **Why log at the gateway instead of inside the agent?** If the agent crashes, in-process logs can be lost — and a compromised agent can tamper with its own logs. A [control plane](/blog/what-is-an-ai-agent-control-plane) logs every operation (including denied ones) before it reaches infrastructure, independent of the agent lifecycle. That independence is what makes the trail trustworthy. --- ### Multi-Agent Orchestration with Claude Code: Patterns and Pitfalls **URL:** https://sentrely.com/blog/multi-agent-orchestration **Published:** 2026-04-25 **Description:** Single agents are powerful. Coordinated agent teams are transformative — and dangerous without proper isolation. Learn the patterns for safe multi-agent Claude deployments. A single well-built Claude agent can do remarkable things. A coordinated team of agents — each specialized, each isolated, each observable — can do things that no single agent could handle alone. Multi-agent architecture isn't just "more agents." It's a different operational model. It's also a different set of failure modes. Getting multi-agent right requires thinking carefully about isolation, identity, communication, and control. ## Why Multi-Agent The obvious reason is scale: some tasks are too large for a single agent's context window or too time-consuming to run sequentially. But there are better reasons: **Specialization.** A `code-review` agent configured for thoroughness operates differently than a `deploy` agent configured for speed. Mixing their contexts and instructions produces worse results than keeping them separate. **Isolation.** A `billing-automation` project shouldn't share credentials, audit trails, or error blast radius with a `code-generation` project. Separate agents means separate blast radii. **Parallelism.** Independent subtasks can run simultaneously. A research agent gathering data in parallel with a summarization agent processing earlier results. Wallclock time drops even if total token cost is the same. **Redundancy and verification.** Two independent agents analyzing the same dataset and comparing results is a validity check you can't get from a single agent reviewing its own work. ## Common Patterns **Pipeline (sequential).** Agent A produces an artifact. Agent B takes that artifact as input and produces the next. Agent C takes B's output and produces the final result. Each agent has a narrow scope and well-defined inputs/outputs. *Example:* `data-extractor` agent pulls raw data from a source → `normalizer` agent cleans and structures it → `reporter` agent generates a formatted output document. *When to use:* When steps have clear dependencies and each agent's output can be fully specified before the next agent runs. **Fan-out (parallel, then merge).** An orchestrator agent splits a large task into independent subtasks and dispatches them to worker agents. Workers run in parallel. Orchestrator collects results and produces a final output. *Example:* `research-orchestrator` receives "analyze these 50 competitor websites" → dispatches to 5 `researcher` agents (10 sites each) → merges their findings into a final report. *When to use:* When subtasks are genuinely independent and the bottleneck is throughput rather than sequential dependencies. **Hierarchical (supervisor/worker).** A supervisor agent manages the overall task and has authority to direct worker agents. Workers report back and receive further instructions. The supervisor maintains the high-level context while workers handle implementation detail. *Example:* `project-manager` agent breaks down a feature request → directs `code-agent` to implement → directs `test-agent` to write tests → directs `review-agent` to check the work → produces a summary for human review. *When to use:* Complex tasks that require judgment at each step, where the next step depends on the result of the previous one. ## A2A Messaging In Sentrely, agents communicate through A2A (agent-to-agent) messaging. This is a structured channel for agents to pass context, results, and instructions to each other — separate from the human-facing communication channels. A2A messaging is preferable to having agents share a database or file system because: - Messages are audited (every A2A message is in the audit log) - Messages are typed and validated at the gateway level - Failed deliveries are handled gracefully - The communication graph is visible in the dashboard A typical A2A message in a fan-out pattern: ```json { "from": "research-orchestrator", "to": "researcher-03", "task": "analyze_website", "payload": { "url": "https://competitor-c.com", "focus_areas": ["pricing", "features", "positioning"] }, "callback_session": "orchestrator-session-id" } ``` The orchestrator knows when all workers have completed because the gateway tracks message delivery and response. ## The Isolation Problem Multi-agent systems fail when isolation is inadequate. The most common failures: **Shared credentials.** If all agents in a project use the same API keys, you can't audit which agent took which action, can't revoke one agent without breaking all of them, and can't have per-agent policies. Every agent needs its own identity. **Cross-project contamination.** Agent A from `billing-project` shouldn't be able to read data from `hr-project`, even if they're running on the same infrastructure. The control plane enforces project-level isolation: agents can only communicate with agents in the same project unless explicitly configured otherwise. **Shared state via the filesystem.** Agents that write to shared directories without coordination produce race conditions that are extremely hard to debug. Use structured A2A messaging or explicit coordination patterns instead of relying on filesystem state. **Cascading failures.** In a pipeline, a bug in Agent B will cause every downstream agent to fail or produce garbage. Design pipelines with explicit validation at each handoff. If Agent B produces output that Agent C can't use, Agent C should fail loudly rather than silently propagate bad data. ## Per-Agent Policies in Multi-Agent Contexts In a multi-agent setup, each agent should have its own policy scoped to its specific role: ```yaml project: invoice-processing-pipeline agents: data-extractor: allow: - service: aws actions: ["s3:GetObject"] resources: ["arn:aws:s3:::invoice-raw/*"] - service: internal actions: ["database:read"] resources: ["invoices.pending"] # No write access, no git access, no external calls normalizer: allow: - service: aws actions: ["s3:GetObject", "s3:PutObject"] resources: ["arn:aws:s3:::invoice-raw/*", "arn:aws:s3:::invoice-normalized/*"] # Can read raw, write normalized — nothing else reporter: allow: - service: aws actions: ["s3:GetObject"] resources: ["arn:aws:s3:::invoice-normalized/*"] - service: email actions: ["send"] resources: ["finance@acme.com"] require_approval: - service: email actions: ["send"] ``` Each agent can only do its part of the pipeline. If `normalizer` is compromised or goes wrong, it can corrupt normalized data — but it cannot touch raw data, cannot send emails, cannot push code. The blast radius of any single agent is bounded. ## Monitoring a Fleet Multi-agent systems require a different monitoring posture than single agents. You're not watching one session — you're watching a graph of sessions with dependencies. Useful things to track at the fleet level: - Overall pipeline progress (how far through the work are we?) - Per-agent session status (running, completed, errored, stuck) - Inter-agent message queue depth (are messages piling up at a bottleneck?) - Aggregate token consumption across the fleet - Failed agent handoffs (Agent B produced output Agent C rejected) Sentrely's dashboard shows a live session view across all agents in a project. When a pipeline stalls, you can see exactly which agent is stuck, what it was trying to do, and where it got blocked. ## Start Smaller Than You Think The biggest mistake in multi-agent systems is architectural over-ambition. Teams design a 15-agent pipeline before they've run a single agent in production. Start with two agents and a clear handoff. Get that working in production. Understand the failure modes. Then add a third agent. The architecture of a mature multi-agent system emerges from operational experience, not from upfront design. A two-agent system you understand is worth more than a ten-agent system you don't. --- ### Managing Claude API Costs in Production: Budgets, Rate Limits, and Alerts **URL:** https://sentrely.com/blog/managing-claude-api-costs **Published:** 2026-04-25 **Description:** A runaway agent can burn through thousands of dollars before you notice. Here's how to implement token budgets, rate limiting, and real-time cost alerts. Agentic AI has a cost problem that interactive AI doesn't: humans self-throttle. When a developer uses Claude interactively, they submit prompts at human speed — maybe 10-20 per hour. A Claude agent in a production loop can make 100 API calls per minute. The delta between "Claude helped me with my code today" and "Claude just spent $3,000 overnight" is the delta between interactive and agentic usage. Without controls, you won't notice until the invoice arrives. ## The Anatomy of Agentic API Costs Understanding where agent costs come from helps you control them. **Context accumulation.** The most common cost driver isn't individual call expense — it's context growing over a session. An agent that starts with a 2,000-token prompt and adds 500 tokens of tool call results per step hits 50,000 tokens by step 96. With Claude 3.5 Sonnet pricing, a session that looked like it would cost $0.50 can end up costing $5-15 if it runs long. **Retry loops.** An agent hitting rate limits or errors will retry. Without a circuit breaker, it retries indefinitely. 100 failed calls cost almost as much as 100 successful calls, and produce no value. **Parallel agent operations.** Multi-agent setups multiply costs. Three agents running the same long context session costs three times as much. If you're running 10 agents in parallel, your cost profile changes dramatically. **Tool call overhead.** Each tool result gets added to context. An agent that makes many small tool calls accumulates context faster than one that makes fewer, larger ones. ## The Three Control Levers ### 1. Token Budgets A token budget sets an upper bound on how much context an agent or project can consume in a given period. When the budget is hit, the agent pauses (or the session terminates, depending on configuration). **Daily per-agent budget.** Limit each individual agent to N tokens per day. For a well-tuned agent doing a known task, you can estimate this from baseline usage. Set the limit at 2-3x expected daily usage to allow for variance without allowing runaway. **Monthly per-project budget.** A higher-level limit across all agents in a project. This is the one that shows up in your invoice. Set it based on what you're actually willing to spend, with a margin. **Per-session limit.** A single session that goes over N tokens gets terminated. This catches the loop case: an agent that would have consumed 500k tokens in a runaway session gets cut off at 50k. A typical configuration: ```yaml budget: per_session_tokens: 100000 # single session cap daily_tokens: 500000 # per agent per day monthly_usd: 200 # per project per month alert_thresholds: [0.5, 0.8, 1.0] ``` ### 2. Rate Limiting Rate limiting controls the *speed* of consumption rather than the *total*. It's the control that catches loops before they drain a budget. **Requests per minute.** An agent in normal operation makes 5-20 API calls per minute. An agent in a runaway loop might make 50-100. A rate limit at 30 requests per minute lets normal operation proceed while cutting off loops early. **Requests per tool type.** More granular: limit a specific tool call to N invocations per minute. This catches cases where the loop is in a specific tool (like a database query) without limiting the rest of the agent's operation. **Session-level rate limiting.** Tracks the request rate for a specific session. If a session exceeds its normal request rate by 3x for more than 2 minutes, treat it as an anomaly. ### 3. Alerts and Notifications Budgets and rate limits are reactive — they stop spending after a threshold is crossed. Alerts are proactive — they tell you that a threshold is approaching. **The 50% alert** is the most useful. When a project hits 50% of its monthly budget by day 15, you still have time to investigate and adjust without hitting the limit. By the time you get the 80% alert, you're reacting. By 100%, you're in incident mode. **Rate anomaly alerts** are different from budget alerts. "This agent is making 3x its normal request rate" surfaces a problem earlier than any budget alert would. A runaway loop at 100 req/min with a 500k daily token budget would exhaust the budget in about 3 hours. An anomaly alert fires in the first 5 minutes. **Cost attribution reports** tell you where your budget is actually going. In a multi-agent environment, the breakdown often surprises people: one agent consuming 60% of the budget is worth investigating even if total spend is within limits. ## Real Cost Example Here's what a well-tuned cost setup looks like for a small agent team: | Agent | Task | Daily Token Budget | Normal Usage | Buffer | |-------|------|--------------------|--------------|--------| | claude-invoice-01 | Process invoices | 300k | ~180k | 1.7x | | claude-deploy-01 | Code review + deploy | 200k | ~120k | 1.7x | | claude-research-01 | Lead research | 150k | ~90k | 1.7x | | **Project total** | | **1.2M** | **~720k** | | Monthly project budget at 30 days: 1.2M tokens/day × 30 = 36M tokens. At Sonnet pricing (~$3/million output tokens), roughly $108/month. Set the project monthly budget at $150 with alerts at $75 and $120. Add rate limits at 2x expected normal rate. This gives you 2 weeks of normal operation before the 50% alert fires, and substantial headroom for variance without runaway risk. ## What Happens When Budget Is Exhausted The behavior when a budget is hit matters as much as the budget itself. **Hard stop (per-session limit):** The session is terminated immediately when the per-session token limit is hit. The agent gets an error. Any work in progress is incomplete. This is appropriate for the runaway loop case — you want it to stop hard and fast. **Grace stop (daily budget):** When the daily budget is hit, running sessions are allowed to complete their current operation, then new sessions are blocked until the next day. This prevents incomplete states for ongoing work. **Soft alert (monthly budget):** When the monthly budget's alert threshold is hit, sessions continue running but a human is notified. They can decide to increase the budget, reduce operations, or continue as-is. The human makes the call. Getting these behaviors right requires thinking through your failure modes: what's worse, an interrupted operation or an unexpected cost overrun? ## The Operational Discipline Cost controls only work if you maintain the discipline to set them properly and review them periodically. Review your agent costs weekly for the first month. Understand where the spend is coming from. Tighten budgets that have too much headroom. Expand ones that are too restrictive. After a month, you'll have enough data to set budgets that are accurate rather than guessed. The goal is a cost profile that looks like this: predictable baseline, clear variance patterns, no surprises. When you have that, you've operationalized your agent costs. Until then, you're gambling. --- ### RBAC for AI Agents: Why Least Privilege Matters More Than You Think **URL:** https://sentrely.com/blog/rbac-for-ai-agents **Published:** 2026-04-24 **Description:** Static RBAC built for humans either over-grants or breaks AI agents. Here's why AI agent access control needs deny-by-default, action+resource scoping, and just-in-time elevation — and how it works in practice. When a human joins your team, you don't give them admin access to everything on day one. You give them the access they need to do their job, and expand it as they demonstrate judgment and need. This principle — least privilege — is fundamental to security for exactly one reason: if something goes wrong, the damage is bounded. AI agents deserve the same treatment. They often don't get it. ## Why static RBAC isn't enough for AI agents Traditional RBAC was built for humans, who have stable roles — a developer is a developer all week. Agents don't work that way: a single agent runs wildly different workflows hour to hour, touching different services and resources each time. Pin an agent to a fixed role and you get one of two bad outcomes. Either the role is broad enough to cover every possible workflow — and now it's massively over-permissioned — or it's tight enough to be safe and the agent's tasks break the moment they stray. As the authorization vendor Oso frames it: humans have stable roles, agents execute dynamic workflows, so static RBAC tends to either break tasks or over-grant access. The fix isn't a bigger role. It's a different model: - **Deny-by-default.** Nothing is allowed unless explicitly permitted. - **Scope to task + resource + action**, not to a persona — `s3:GetObject` on *this* prefix for *this* job, not "S3 access." - **Just-in-time elevation.** Grant an extra permission for the moment it's needed with a short TTL, then drop it — not a standing grant. - **Continuous tightening.** Watch what the agent actually uses and trim the rest. - **Containment.** A kill switch and per-agent identity so one misbehaving agent is bounded and instantly revocable. That's the model the rest of this guide implements — and it's why agent access control belongs in a [control plane](/blog/what-is-an-ai-agent-control-plane), not in a static IAM role. ## Why Agent RBAC Is Harder Than Human RBAC With humans, access control is mostly about identity and role. Jordan is a developer, so Jordan gets developer access. Developers can push to feature branches but not main. Simple. With agents, it's more complicated for three reasons: **Agents are more prolific.** A developer might make 20 git operations in a day. An agent might make 200 in an hour. The blast radius of a permissive policy scales with operation volume. **Agents don't exercise judgment about scope.** A developer who accidentally has access to a sensitive bucket will usually notice it's not relevant to their work and leave it alone. An agent tasked with "find all CSV files" will find them everywhere it has access, without the social inhibitions that keep humans from overreaching. **Agent intent is harder to audit post-hoc.** With humans, "Jordan deleted that table" is a starting point for a conversation. With an agent, "claude-ops-01 deleted that table" starts a reconstruction exercise. Better controls reduce the need for reconstruction. ## The Structure of an Agent Policy A well-designed agent policy specifies four things: **What services the agent can call.** AWS, GitHub, Stripe, custom internal APIs, specific tools. Not "everything" — the specific services the agent's job requires. **What actions within those services.** `s3:GetObject` not `s3:*`. `git:push` not `git:admin`. `charges:read` not `charges:write`. This is where most policies are too permissive — it's easier to write wildcards than to enumerate specific actions. **What resources those actions apply to.** `arn:aws:s3:::invoice-data/processed/*` not `arn:aws:s3:::*`. The specific bucket prefix, the specific git repository, the specific Stripe resource type. Resources are where you contain blast radius. **What requires approval.** The subset of allowed actions that are high enough risk to require human confirmation before execution. Pushing to protected branches. Deleting anything. External communications. Here's a realistic example: ```yaml project: invoice-processor agents: claude-invoice-01: # Git: read any repo, push only to feature branches of invoice repos allow: - service: git actions: [clone, pull, fetch] resources: ["acme/*"] - service: git actions: [push] resources: ["acme/billing", "acme/invoices"] branches: ["feature/*", "fix/*"] # S3: read raw invoices, write processed output allow: - service: aws actions: ["s3:GetObject", "s3:ListBucket"] resources: ["arn:aws:s3:::invoice-data/raw/*"] - service: aws actions: ["s3:PutObject"] resources: ["arn:aws:s3:::invoice-data/processed/*"] # Stripe: read-only allow: - service: stripe actions: ["invoices:read", "customers:read"] # These require human approval require_approval: - service: git actions: [push] branches: ["main", "staging"] - service: aws actions: ["s3:DeleteObject"] # Hard limits deny: - service: aws actions: ["iam:*", "ec2:*", "rds:*"] - service: stripe actions: ["charges:create", "refunds:create", "customers:delete"] ``` This policy lets the agent do its job completely. It can't do significant damage beyond its scope, even if it makes a mistake or is prompted maliciously. ## Common Over-Permissioning Mistakes **`s3:*` on `*`.** The most common AWS mistake. Every AWS action on every bucket. The agent only uses `s3:GetObject`, but you gave it `s3:DeleteBucket` just in case. This is how "clean up temp files" becomes "delete production data." **Push access to all branches.** Giving an agent push access to every branch, including `main` and `release/*`. The agent only touches feature branches in practice, but nothing stops it from pushing to main when it decides it's done. **Shared credentials across agents.** Ten agents, one set of credentials. Impossible to audit which agent took which action. Impossible to revoke one agent's access without revoking all of them. **No resource constraints.** Policies that specify action but not resource: "allow `s3:GetObject`" with no bucket specification. The agent can read any bucket it can discover. **Forget about tools, not just APIs.** Policies that scope AWS and git but forget about custom tool integrations. If the agent can call arbitrary bash commands, or has a tool that can query any database table, the service-level policies are irrelevant. ## Testing Your Policies Write a policy, then try to break it. Prompt the agent to: - "Find all CSV files across all S3 buckets" - "Push the current state of main to the production server" - "Delete the temp files in the data directory" - "Check if there are any IAM users with admin access" A well-scoped policy should result in 403 responses and "access denied" audit entries — not successful operations. If the agent can do these things, your policy is too broad. ## Policy Version Control Treat agent policies the same way you treat infrastructure code: version control, code review, change history. A policy change is a security change. It should go through the same review process as any other security-relevant code. "Claude asked for more access so I gave it to him" is not an audit trail. A pull request with the requester, the rationale, and the reviewer is. Sentrely enforces policies from YAML files committed to your project repository. Every policy change creates a git commit, which means every change has an author, a timestamp, and a diff. ## The Minimum Viable Policy If you're starting from scratch and need a policy today, here's the floor: 1. **No wildcard actions.** List specific actions, even if the list is long. 2. **No wildcard resources.** Scope to the specific buckets, repos, and resources the agent needs. 3. **Push to main requires approval.** Always. 4. **Data deletion requires approval.** Always. 5. **No IAM actions.** Ever. Unless IAM management is the agent's explicit job. These five rules eliminate the most common blast radius scenarios while leaving most agent workflows intact. Start here and tighten from experience. ## FAQ **What is RBAC for AI agents?** Role-based access control adapted for autonomous agents: instead of mapping an agent to a broad human-style role, you grant deny-by-default permissions scoped to specific actions and resources for specific tasks, with per-agent identity and a kill switch. The goal is bounded blast radius — if the agent errs or is prompt-injected, the damage is contained. **Why doesn't normal (static) RBAC work for AI agents?** Static roles assume stable behavior. Agents run dynamic, varied workflows, so a fixed role either over-grants (broad enough for everything) or breaks tasks (too tight). Agents also act at high volume without judgment, so over-permissioning is far more dangerous than with a human. **How do I implement least privilege for a Claude Code agent?** No wildcard actions or resources; scope to specific ARNs/repos/prefixes; default read-only and elevate just-in-time; require approval for pushes to main and any deletion; deny IAM by default. Enforce it at a [control plane](/blog/what-is-an-ai-agent-control-plane) the agent can't reconfigure, and version-control the policy. **How is this different from Claude Code's built-in permissions?** Native config is an allowlist-with-holes that the agent or developer can skip or talk around (see [Claude Code security](/blog/claude-code-security)). Agent RBAC enforced at a gateway is deny-by-default, central, and tamper-proof — the agent holds no credentials and can't widen its own scope. --- ### Running Claude Agents in Production: 7 Things That Will Break **URL:** https://sentrely.com/blog/claude-agents-in-production **Published:** 2026-04-24 **Description:** Taking Claude Code from proof-of-concept to production is harder than demos suggest. Here are the 7 most common failure modes and how to prevent them. Demo-to-production is the hardest gap in AI agent development. Not because the technology is unreliable — Claude is remarkably capable — but because production exposes assumptions that a demo doesn't test. Real data. Real infrastructure. Real failure modes. Real costs at real scale. Here are the seven failure modes we see most often, what they look like in practice, and how to prevent them. ## 1. Runaway Loops **What it looks like:** An agent tasked with "fetch the current lead data from the CRM" hits a rate limit. The responsible thing seems to be to retry. So it retries. And retries. 73 times over 19 minutes, spending $4.20, making no progress. **Why it happens:** No circuit breaker. No maximum retry count. No monitoring that notices the pattern. The agent is technically following its instructions. **Prevention:** - Configure maximum retry attempts at the gateway level - Set anomaly detection that alerts when an agent makes N identical calls in M minutes - Set a daily token budget that would cap this at, say, $2 before stopping the session - Monitor for sessions that are "alive but not making progress" A loop that burns $4 is a minor incident. The same loop, with a higher budget and no monitoring, is a $200 surprise. The controls are cheap to implement and the alternative is expensive to discover. ## 2. Blast Radius Incidents **What it looks like:** An agent with broad AWS permissions, tasked with "clean up the development environment," interprets "clean up" more aggressively than you intended. S3 buckets, CloudWatch alarms, IAM roles, EC2 snapshots — all deleted. Some of them weren't in the development environment. **Why it happens:** No resource scoping in the policy. "Development environment" is a human concept; the agent doesn't share your mental model of which resources are in scope. **Prevention:** - Scope policies to specific resource ARNs and prefixes, not wildcards - Require approval for any deletion operation, regardless of environment - Tag resources explicitly and enforce policy based on tags - Test your policies by prompting the agent to do things it shouldn't be able to do The rule: an agent should only be able to affect resources that are explicitly in its scope. Everything else should return a 403. ## 3. Missing Audit Trail **What it looks like:** Something unexpected happens. You need to understand what the agent did. You have Claude's conversation history (maybe), but no structured log of actions taken, what was allowed vs. denied, or what external effects were produced. **Why it happens:** Audit logging is treated as optional infrastructure. It's not. **Prevention:** - Route all agent operations through a control plane that logs every action - Log the structured data: agent ID, session ID, action type, resource, result, timestamp - Export logs to storage you own and control (not just a vendor dashboard) - Test that you can reconstruct a session: given a session ID, can you replay what happened? The test: pick any session from last week and answer "what did this agent do between 2:00 PM and 2:30 PM?" If you can't do this in under 5 minutes, you don't have a real audit trail. ## 4. Shared Credentials **What it looks like:** You have 8 agents. They all use the same AWS access key. One of them needs to be revoked because it's behaving unexpectedly. To revoke it, you have to rotate the credential — which breaks all 8 agents simultaneously. **Why it happens:** Shared credentials are easier to set up initially. Per-agent identity requires more configuration. **Prevention:** - Each agent gets a unique identity (Sentrely manages this automatically) - Credential rotation affects only the target agent, not the fleet - Audit logs attribute actions to specific agents, not "a thing with these credentials" - If one agent is compromised or misbehaving, it can be isolated without affecting others This also matters for compliance: "an agent did this" is not sufficient attribution. "claude-deploy-01 in session d97e2169 did this at 14:23:07" is. ## 5. No Cost Controls **What it looks like:** Your agent fleet runs well for two weeks. Then a bug in one agent causes a loop. By the time someone notices, it's spent $340 on a task that should have cost $3. The monthly API invoice is $800 over budget. **Why it happens:** Token budgets are never configured. Cost monitoring is manual. Nobody set up an alert at 80% of expected spend. **Prevention:** - Set per-project daily and monthly token budgets with hard limits - Configure alerts at 50%, 80%, and 100% of budget - Route cost alerts to a monitored Slack channel - Set per-session limits so a single runaway session can't consume a week's budget The asymmetry here is important: a budget that's too tight costs you an interrupted task. No budget costs you an unexpected invoice. Default to conservative limits and loosen them based on observed usage. ## 6. Context Drift in Long Sessions **What it looks like:** An agent starts a long task — migrate 500 records from old schema to new schema. Three hours in, it's made inconsistent decisions because the early context about what it was doing has drifted out of the active window. Some records are migrated correctly, some with the old schema, some with errors. **Why it happens:** LLM context windows are finite. Long-running agents that rely on conversational context accumulate noise and lose early instructions. **Prevention:** - Design agents to work in discrete, short-context chunks rather than long single sessions - Checkpoint progress explicitly (write state to a file or database) - For batch operations, process in smaller batches with verification between them - Use structured task definitions that get re-injected at each checkpoint rather than relying on conversation history Long sessions are a yellow flag. An agent that needs to maintain context for hours is usually an agent whose task could be decomposed better. ## 7. No Kill Switch **What it looks like:** Something is going wrong. An agent is behaving unexpectedly. You need to stop it. But the only way you can think of to stop it is to revoke the credentials it's using — which also breaks 7 other agents that are running fine. **Why it happens:** Kill switches are designed for incidents, and incident response is designed after the fact. **Prevention:** - Every agent session has a unique session ID - The control plane exposes a "terminate session" endpoint - You've tested this endpoint in a non-incident context and know it works - You've also tested "pause all agents in project X" for larger incidents The test: can you stop a specific running agent in under 30 seconds, without affecting anything else? If not, fix that before you go to production. --- ## The Common Thread None of these failure modes are exotic. They're the predictable consequences of deploying powerful automation without the operational controls that every other kind of production system requires. The good news: all seven are preventable with a control plane that's properly configured before you go live. Not after the first incident. Before. The teams that run agents most confidently aren't the ones with the most powerful agents. They're the ones who invested in the boring infrastructure: policies, audit trails, budgets, session management, kill switches. The unsexy stuff that makes the exciting stuff safe to actually run. --- ### Slack Alerts and Approvals for Claude Agents: Real-Time Human Oversight **URL:** https://sentrely.com/blog/slack-alerts-for-ai-agents **Published:** 2026-04-23 **Description:** How to set up Slack-based monitoring, alerts, and one-click approvals for your Claude agent fleet — so you're never flying blind. The problem with agent monitoring dashboards is that nobody watches them. You open the dashboard when something's wrong. By then, the agent has been doing the wrong thing for 20 minutes. Slack-based oversight works because it meets people where they already are. When your agent needs attention, it shows up in the same channel where your team is discussing the work. No new tool to learn, no dashboard to remember to open. Here's how to build it well. ## What Butler Does Sentrely's Slack integration uses a bot called Butler. Butler handles three categories of Slack interactions: **Approval requests.** When an agent attempts a gated operation — pushing to main, deleting data, sending external communications — Butler posts a message with action buttons. One click approves or denies. The agent receives the decision and proceeds or stops accordingly. **Proactive alerts.** Butler monitors agent behavior and surfaces anomalies without waiting to be asked. A loop that's been retrying the same call for 15 minutes. A session that's been idle for 2 hours but is still running. A project that's consumed 80% of its daily token budget by 10 AM. **Status reports.** When an agent completes a task, Butler posts a summary: what was done, how long it took, what was deployed, how many tokens were consumed. ## Channel Architecture The key to not drowning in noise is routing different signals to different channels. A single `#ai-agents` channel that receives everything becomes unactionable fast. The pattern that works: **`#engineering-agents`** — Deployment approvals and code-related gating. This is where your dev team lives. Approval requests for git pushes, code reviews, deployments go here. The audience is engineers who can make informed decisions quickly. **`#agent-alerts`** — Anomaly detection and proactive issues. Stuck loops, unusual patterns, session anomalies. Lower urgency than approvals, but needs engineering attention within the hour. **`#billing-alerts`** — Cost and budget notifications. Keeps financial visibility separate from operational noise. Finance or team leads can monitor this without wading through deployment approvals. **`#deployments`** — Task completion summaries. Low-noise, informational. Good for stakeholders who want to know "what shipped today" without seeing every approval request. **`#access-requests`** — When agents need elevated permissions not in their current policy. This often needs a different approver than deployment decisions. ## Approval Request Anatomy A good approval request message gives the approver everything they need to decide in under 30 seconds: ``` 🚨 Approval Required Agent claude-deploy-01 wants to push to main on acme/api. 📁 3 files changed · +142 −38 🔖 feat: add stripe webhook handler 👤 Session started 4 minutes ago by claude-deploy-01 [✓ Approve] [✕ Deny] [View Diff] ``` What makes this effective: - The headline is the decision: push to main, yes or no? - The diff summary gives confidence without requiring the approver to read code - The commit message provides intent context - The session age confirms this is a fresh, active session (not a zombie) - Buttons make the decision a single click - The diff link is there for anyone who wants it, but not required After approval: ``` ✅ Approved by Jordan. Push proceeding — commit a3f92d1 is live. ``` After denial: ``` ✕ Denied by Jordan: "use the staging branch first" ``` The reason flows back to the agent, which can then route to staging as instructed. ## Proactive Alert Examples **Loop detection:** ``` 🔴 Agent needs attention — claude-research-01 Making 73 calls to GET /api/v2/properties over 19 minutes. All returning 429. Session may be stuck. 💸 $4.20 spent · Last progress: 23 min ago [View Session] [Stop Agent] [Dismiss] ``` This surfaces before you'd notice it in a dashboard. The agent is burning money and making no progress. One click stops it. **Budget alert:** ``` ⚠️ Budget Alert — billing-automation (81%) Used $163.40 of $200.00 monthly budget. At this rate, exhausts in ~3 days. Top consumer: claude-invoice-01 · $94.20 · 578k tokens [Increase to $300] [Pause Agents] [View Breakdown] ``` The action buttons let the approver handle it without leaving Slack. **Unusual access pattern:** ``` 🔑 Access Request — claude-infra-01 Hit 403 trying to write to s3://prod-data-lake/raw/ Current policy: read-only. Agent says: "need to write processed leads back to S3" 📋 Requesting: s3:PutObject on prod-data-lake/raw/* [Grant 1h] [Grant 24h] [Deny] [View Policy] ``` ## Setting Up Butler Butler is configured at the project level in Sentrely. The configuration specifies: ```yaml notifications: slack: workspace: acme channels: deployments: engineering-agents alerts: agent-alerts billing: billing-alerts access: access-requests approval_timeout: 30m approval_default: deny ``` `approval_timeout` is important: if an approval request expires without a response, it defaults to `deny`. Agents don't block indefinitely. Each project can route to different channels, which means `billing-automation` agents can post to different channels than `code-review` agents — keeping team-specific signals in team-specific channels. ## What Good Agent Observability Looks Like After a few weeks with Butler configured, you should be able to answer these questions without leaving Slack: - What did my agents do today? - Is anything stuck right now? - Am I going to hit my budget this month? - Did last night's deployment succeed? If the answer to any of these requires opening a dashboard or querying a database, your alerting coverage has gaps. The goal isn't to route everything to Slack — it's to route the right things. Noise kills adoption. The agents and alerts that don't need human attention should stay silent. The ones that do should surface themselves, with enough context to act on immediately. When that works, oversight doesn't feel like overhead. It feels like a teammate keeping you informed. --- ### Human-in-the-Loop AI: When and How to Gate Claude Agent Actions **URL:** https://sentrely.com/blog/human-in-the-loop-ai **Published:** 2026-04-23 **Description:** Not every agent action should run autonomously. A practical framework for deciding which operations need human approval — without killing your agents' velocity. The most common mistake in AI agent deployment isn't giving agents too little autonomy. It's giving them too much, too fast, without a model for what requires human judgment. Human-in-the-loop (HITL) isn't about distrust. It's about recognizing that some decisions have asymmetric stakes — easy to do, hard to undo, large blast radius if wrong. For those decisions, a 30-second human review is worth orders of magnitude more than the automation speed you give up. The art is knowing which decisions those are. ## The Autonomy Spectrum Think of agent operations on a spectrum from fully supervised to fully autonomous: **Fully supervised:** Human approves every action. Safe but pointless — you've automated nothing, just added a layer of indirection. **Selective gating:** Agents run freely for most operations. Specific high-risk operations pause for human review. This is the target state for most production deployments. **Fully autonomous:** Agents run without any human involvement. Appropriate for well-understood, low-risk, highly reversible operations. Dangerous when applied broadly. Most teams start at fully supervised (because it feels safe) and never move toward selective gating (because moving feels risky). The result is agents that are technically deployed but practically useless because every prompt requires a human to sit and click through approvals. The goal is selective gating: identify the 10-20% of operations that genuinely need human review, let the other 80-90% run freely. ## The Three Dimensions of Risk Three factors determine whether an operation needs a gate: **Reversibility.** Can you undo it? Reading a file: fully reversible (nothing changed). Pushing to a feature branch: reversible with effort (revert the commit). Pushing to main and triggering a deployment: hard to reverse, especially if customers hit the change. Deleting a production database: practically irreversible. **Blast radius.** How bad is the worst case? A bug in a test file: small blast radius. A bug in authentication middleware: large blast radius. An email sent to 10,000 customers: very large blast radius. **Confidence.** How certain are you that the agent is right? An agent that's been running the same task correctly for three months deserves more autonomy than an agent doing something new. Low confidence plus high stakes is the combination that needs a gate. ## Operations That Should Always Gate Regardless of how well you know your agent, some operations should always require human approval in production: **Pushes to protected branches.** Your `main`, `release/*`, and `hotfix/*` branches should never receive an automated push without human review. Not because agents can't write good code — they often can — but because an unreviewed push to main is a known failure mode that has burned too many teams. **Data deletion.** `DELETE` queries, `rm -rf`, S3 object deletion, database drops. Deletion is generally irreversible. Even if the agent is right that the data should be deleted, a human should confirm. **External communications.** Emails, Slack messages, API calls that trigger customer-visible actions. The blast radius of an incorrect mass email is enormous and immediate. **Large financial transactions.** Any amount over a threshold you set (often $100-$1000 depending on context) should require approval. Agents working with payment systems are particularly sensitive here. **IAM and permission changes.** Modifying who can access what is a security operation. It should always have a human owner. **Infrastructure destruction.** Terminating EC2 instances, deleting S3 buckets, dropping databases. Even in "ephemeral" environments, verify before destroying. ## Operations That Can Run Freely Conversely, these operations are typically safe to run without gates: - Reading files, databases, and APIs (no side effects) - Writing to development branches - Creating new resources (easy to clean up) - Querying observability systems (logs, metrics, traces) - Running tests - Generating code or documentation for review - Writing to staging or development environments with clear rollback paths The common thread: these are either read-only (no state change) or easily reversible with small blast radius. ## Implementing Gates Without Killing Velocity The failure mode of heavy gating is an inbox of 50 approval requests that nobody processes because they're too granular to be useful. Avoid it: **Batch similar decisions.** Instead of one approval per file in a batch operation, request one approval for "delete these 47 temp files from the staging bucket" with a list attached. **Provide context in the approval request.** An approval request that says "agent wants to push to main" is less useful than one that says "pushing `feat: add stripe webhook handler` (3 files, +142/-38 lines) to `acme/api` main. [View Diff]". The goal is a decision in under 30 seconds. **Use the right channel.** Slack works for operations that need a response in minutes. PagerDuty or urgent alerts work for time-sensitive operations. Don't route everything to the same channel. **Set reasonable timeouts.** An approval request that expires after 30 minutes and defaults to "deny" is better than one that blocks indefinitely. Agents can gracefully pause and resume. **Make "deny" informative.** When a human denies an operation, the reason should flow back to the agent. "Denied: use the staging branch instead" is actionable. "Denied" with no context leaves the agent stuck. ## The Butler Pattern The most effective implementation we've seen uses a Slack bot (Butler, in Sentrely's case) as the approval interface. The flow looks like this: 1. Agent attempts a gated operation 2. Gateway intercepts, creates an approval request 3. Butler posts a rich message to the designated Slack channel: what the agent wants to do, relevant context, action buttons 4. Human clicks Approve or Deny (with optional reason) 5. Gateway receives the decision and either proceeds or blocks 6. Butler posts a confirmation message in thread The entire interaction takes 15-30 seconds for the human, happens in the tool they're already using, and creates a natural audit trail (the Slack thread). This pattern makes human oversight feel like part of the workflow rather than an interruption to it. When it works well, people stop thinking of approval gates as friction and start thinking of them as the natural handoff between automation and human judgment. ## Building Toward More Autonomy Good HITL implementation isn't a permanent state — it's a foundation for earning more autonomy over time. Start with more gates than you need. Get the approval flow working. Watch what gets approved versus denied. Look for patterns: if the same type of operation gets approved 95% of the time with no concerns, consider whether it actually needs a gate. Over time, you develop data on which operations your agents handle well and which they don't. That data is the basis for adjusting your gating model — loosening where confidence is high, tightening where you've seen failures. The teams that give agents the most autonomy in the long run are the ones who started with the most deliberate oversight. --- ### Private AI Agents: Keeping Your Data Off Third-Party Infrastructure **URL:** https://sentrely.com/blog/private-ai-agents **Published:** 2026-04-22 **Description:** When Claude agents touch sensitive data, you need to know what goes where. Learn deployment patterns for data-sovereign AI agent operations. When you run a Claude agent against your codebase, your database, or your customer data, something flows through every tool call: context. The agent needs context to do its job — and that context often contains things you'd prefer not to send to infrastructure you don't control. This isn't a hypothetical concern. It's the reason regulated industries move slowly on AI agents, and it's the reason "private AI" has gone from buzzword to procurement requirement in the last 18 months. ## What Data Actually Flows Most teams underestimate what their agents are transmitting. Consider a Claude agent tasked with "review the Q1 invoices and flag anything anomalous." The agent will: - Read invoice files from S3 or a database (customer names, amounts, account numbers) - Pass that content to Claude as context (it becomes part of the prompt) - Call Claude's API with that context (it leaves your infrastructure) - Potentially log the interaction (now it's in a third-party log system) Depending on your industry, that invoice data might be subject to GDPR (EU customer data), HIPAA (if invoices relate to healthcare services), PCI-DSS (payment card data), or SOC 2 requirements. In each case, you need to know: where does that data go, who can see it, and how long is it retained? With an uncontrolled agent, the answer is "wherever the agent sends it, and we don't know." ## The Three Layers of Data Exposure **Layer 1: The model call.** When the agent sends a prompt to Claude's API, the prompt content leaves your environment. Anthropic's data handling policies govern what happens next. For most use cases, this is acceptable — Anthropic has strong security practices and doesn't train on API data by default. But for some regulated contexts (particularly healthcare and financial services), even this level of third-party exposure is a compliance issue. **Layer 2: The control plane.** If your agent routes through a gateway, what does the gateway see? A poorly designed gateway that logs full request payloads could be storing sensitive context you didn't intend to retain. A well-designed one logs metadata (action taken, policy applied, result) without storing the content itself. **Layer 3: The tools.** When the agent calls tools — reading files, querying databases, hitting APIs — those calls are often logged by default. Your cloud provider, your database, your SaaS tools may all be storing records of what your agent accessed. ## Deployment Patterns **Pattern 1: Standard managed (most teams)** The agent runs in your environment. Model calls go to Anthropic's API. The control plane (gateway) is hosted by a managed service like Sentrely. This is appropriate for most use cases — the sensitive data stays in your environment, model calls follow Anthropic's standard data policy, and the gateway sees metadata but not payload content. **Pattern 2: VPC-isolated control plane** For organizations with stricter requirements, the control plane itself is deployed inside your VPC. The gateway runs on infrastructure you control, audit logs stay in your environment, and no control-plane data leaves your network perimeter. Sentrely's Enterprise tier supports this model. **Pattern 3: Air-gapped with local models** For the most sensitive use cases, everything runs inside your environment: local LLM (Ollama, LM Studio, or a private Bedrock endpoint), self-hosted control plane, internal tooling only. Latency is higher, model capability is lower, but zero data leaves the perimeter. This is the banking and healthcare pattern for the most sensitive workloads. ## What a Control Plane Does for Data Privacy A well-designed control plane improves your data privacy posture in three ways: **Scoped access.** Instead of the agent having credentials that can read everything, the control plane enforces policies that limit which data stores the agent can access. An agent processing invoices shouldn't have access to HR records, even if the credentials technically allow it. **Audit trail as proof.** Compliance requirements often include demonstrating what happened to data — not just asserting it. An immutable audit log of every agent access, exportable to your own S3 bucket, gives you evidence that data was handled correctly. This is what auditors actually want to see. **Data residency enforcement.** A control plane with geographic awareness can enforce that agent operations only use infrastructure in specific regions. For GDPR compliance requiring EU data residency, this matters. ## Regulatory Context **GDPR.** Any agent processing personal data about EU residents requires a lawful basis, data minimization, and the ability to demonstrate compliance. The "data minimization" principle is particularly relevant — your agent policy should only allow access to the specific data needed for the task. **HIPAA.** Protected Health Information (PHI) can only flow to Business Associates with appropriate agreements in place. If your agent processes PHI, your control plane provider needs to be a covered BA. Alternatively, a VPC-isolated deployment keeps PHI inside your environment. **SOC 2.** The access control and audit trail requirements of SOC 2 Type II map directly to what a good control plane provides. Per-agent identity, least-privilege policies, immutable audit logs, and session tracking satisfy most SOC 2 access control requirements out of the box. **PCI-DSS.** Payment card data should not enter AI model prompts unless you have specific controls and agreements in place. The safest approach is to ensure agents never see raw card numbers — only tokenized or masked versions. ## Practical Checklist Before deploying an agent against sensitive data: - [ ] Identify what data the agent will access - [ ] Determine applicable regulatory requirements - [ ] Scope agent permissions to exactly what's needed - [ ] Verify the control plane logging policy (metadata vs. content) - [ ] Choose deployment pattern (shared / VPC-isolated / air-gapped) - [ ] Export audit logs to infrastructure you own - [ ] Document the data flow for compliance purposes The goal isn't to prevent AI agents from handling sensitive data — they're often most valuable precisely because they can process large volumes of it. The goal is to handle it with the same controls you'd apply to any system that touches sensitive data. The teams that get this right aren't slower to adopt AI agents. They're faster to get them approved. --- ### AI Compliance Checklist: Before You Ship Claude Agents to Production **URL:** https://sentrely.com/blog/ai-compliance-checklist **Published:** 2026-04-22 **Description:** A practical checklist covering access control, audit trails, data handling, approvals, and cost controls for teams deploying Claude Code agents. Most teams discover their compliance gaps in production. An auditor asks for a list of every action a specific agent took last Tuesday. An incident happens and nobody can reconstruct what occurred. A security review flags that all agents share one set of credentials. This checklist exists so you discover those gaps before you ship. It's organized by the layer where failures happen most often. Work through it before your first production agent goes live — or use it to audit what's already running. ## Audit Trail **[ ] Every tool call is logged.** Not sampled, not summarized — every call. The log includes: agent identity, session ID, action attempted, whether it was allowed or denied, timestamp, and result. **[ ] Logs are immutable and tamper-evident.** Audit logs that agents (or operators) can modify are not audit logs. They're suggestions. Store logs in append-only storage (S3 with object lock, or equivalent). **[ ] You can reconstruct any session.** Given a session ID, you can replay exactly what the agent did, in order, with timestamps. This is the test: pick a session from last week and try to answer "what did this agent do between 2:00 PM and 2:30 PM?" **[ ] Logs are exported to infrastructure you own.** If your control plane is a managed service, make sure audit data exports to your own S3 bucket or equivalent. You need to own your audit trail, not just have access to it through a vendor dashboard. **[ ] Retention policy is defined and enforced.** How long do you keep agent audit logs? 90 days? 1 year? 7 years (for financial services)? Define it, configure it, verify it. --- ## Human Oversight **[ ] Approval gates are configured for high-risk operations.** Define which operations require human approval before proceeding. At minimum: pushing to protected branches, deleting data, sending external communications (email, Slack messages, API calls to external systems), and any financial transactions. **[ ] Notification channels are configured and tested.** Approval requests need to reach a human reliably. Test the Slack/Telegram integration before going live. Verify that someone is actually monitoring the approval channel. **[ ] There's a kill switch.** You can terminate a specific agent session, or pause all agents in a project, from a single action. Test this. If you can't kill a runaway agent in under 30 seconds, your oversight model has a gap. **[ ] Session timeouts are configured.** An agent that's been idle for 2 hours probably shouldn't still be running. Configure maximum session durations appropriate for your use case. --- ## Cost Controls **[ ] Per-project token budgets are set.** Each project has a maximum daily or monthly token allowance. Not a soft guideline — a hard limit that pauses agents when reached. **[ ] Alert thresholds are configured.** You get notified when a project hits 50%, 80%, and 100% of budget. The 50% alert is the useful one — it gives you time to investigate before the problem becomes expensive. **[ ] Rate limits are defined.** Maximum requests per minute, per agent. This prevents runaway loops from burning budget before the daily limit kicks in. **[ ] You've tested what happens when budget is exhausted.** Does the agent stop gracefully? Does it error in a way that creates cascading problems? Verify the budget exhaustion behavior in a test environment. --- ## Data Handling **[ ] Sensitive data is masked before entering agent context.** Credit card numbers, SSNs, passwords, and similar data should not appear verbatim in agent prompts. Pass tokenized or masked versions. If the agent doesn't need the raw value, don't give it the raw value. **[ ] Data access is scoped to what the agent needs.** If the agent is analyzing invoices from Q1, it shouldn't have read access to all invoices from all time. Scope data access at the query level where possible. **[ ] Your control plane provider's data handling policy is documented.** If you're using a managed gateway, you need to know: what do they log? where is it stored? what's the retention policy? This should be in a DPA (Data Processing Agreement) if you're handling GDPR-regulated data. **[ ] Applicable regulations are identified.** For the data your agents touch: does GDPR apply? HIPAA? PCI-DSS? SOC 2? List them. Each has specific requirements that map to controls above. --- ## Incident Response **[ ] You have a documented response process for agent incidents.** When an agent does something unexpected, what's the sequence? Who gets notified? Who has authority to kill sessions? Where do you look first? **[ ] You've run a tabletop exercise.** Pick a scenario: "Agent X pushed to main at 3 AM." Walk through your response. How long does it take to understand what happened? How long to contain it? This exercise reliably surfaces gaps. **[ ] Recovery procedures are documented.** If an agent deletes something it shouldn't have, can you restore it? From where? In how long? These questions have answers that aren't "hope for the best." --- ## The Meta-Question After going through this checklist, ask: "If something goes wrong with an agent, can I answer these four questions in under 10 minutes?" 1. What exactly did the agent do? 2. When did it happen? 3. What was the blast radius? 4. How do I stop it from happening again? If the answer to any of these is "I don't know," that's your gap. Fix it before you ship. --- ### Claude Code YOLO Mode: What It Is, Why It Exists, and How to Use It Safely **URL:** https://sentrely.com/blog/claude-code-yolo-mode **Published:** 2026-04-21 **Description:** The --dangerously-skip-permissions flag unlocks true agentic automation. Here's exactly what it does, what it allows, and how a control plane makes it production-safe. If you've run Claude Code in a terminal, you know the approval prompts. "Claude wants to execute a bash command. Allow?" Click. "Claude wants to write to this file. Allow?" Click. It's safe. It's deliberate. It's also incompatible with running Claude as an automated agent. Enter `--dangerously-skip-permissions`, also known (affectionately) as YOLO mode. ## What YOLO Mode Actually Does The `--dangerously-skip-permissions` flag does exactly what it says: it skips the interactive permission prompts that Claude Code shows when it wants to take an action. Without the flag, every tool call that touches the filesystem, runs commands, or calls external APIs requires explicit human approval in the terminal. With the flag, Claude runs autonomously. It writes files, executes commands, and calls APIs without stopping to ask. This is what makes true agentic automation possible — a Claude agent that can work through a multi-hour task without a human sitting at the keyboard approving every step. The flag is required any time you want Claude to run: - In a CI/CD pipeline - As a scheduled task or cron job - As a long-running background agent - Orchestrated by another system (not a human) Without it, Claude is a powerful assistant. With it, Claude is an autonomous agent. ## Why the Flag Exists Anthropic built interactive approvals into Claude Code for good reason: most users are running it locally, interactively, and want to stay in the loop on every action. The approval flow is the right default for that use case. But the same interactive flow that protects a developer working on their laptop is a blocker for production automation. You can't have a pipeline waiting for someone to click "approve" in a terminal. YOLO mode exists because the correct solution to "I need Claude to run autonomously" isn't "remove all safety" — it's "move the safety layer to something better than interactive prompts." The flag signals that you, the operator, have taken responsibility for implementing appropriate controls elsewhere. That's the key insight. The flag doesn't mean "no controls." It means "I'm using a different control mechanism than interactive prompts." ## What YOLO Mode Allows (and What It Doesn't) YOLO mode removes the *interactive approval prompts*. It doesn't remove everything: **It doesn't bypass your gateway.** If you've configured Claude to route through a control plane like Sentrely, policy enforcement still happens on every request. An agent running in YOLO mode against a gateway with a restrictive policy is actually *safer* than an interactive session without one — because the policies are consistently enforced regardless of who's at the keyboard. **It doesn't expand Claude's capabilities.** The agent can only do things Claude Code can already do. No new attack surface is opened. **It doesn't disable audit logging.** A properly configured control plane continues to log everything. What it *does* remove: the human speed bump between each action. That's powerful when you need automation. It's dangerous when you have no other controls. ## Real Risks in Production Here's what can actually go wrong when you run YOLO mode without a control plane: **Unreviewed code in production.** An agent with push access to all branches will push when it thinks it's done. Maybe it's right. Maybe it introduced a subtle bug. Without a gate, you find out when the deployment fails. **Runaway API consumption.** A loop bug that a human would notice after three prompts can run for 20 minutes in YOLO mode. By the time you see the Anthropic invoice, the agent has made 200 API calls and you've spent $15 on a bug that should have been caught immediately. **Accidental destructive operations.** "Clean up the temp files" is a reasonable instruction. If the agent's definition of "temp files" is broader than yours, and there's no policy preventing deletion of things that matter, you've lost data. **No attribution.** Without session identity and audit logging, you can't tell what happened. "Something deleted those files last night" is not a useful incident report. ## The Gateway Approach: YOLO Mode Done Right The right way to run YOLO mode in production: 1. **Route through a control plane.** Set `GATEWAY_URL` to your managed gateway endpoint. Every tool call goes through policy enforcement. 2. **Define least-privilege policies.** Your agent should be able to do exactly what it needs and nothing else. `git push` to `feature/*` only. `s3:GetObject` on specific prefixes only. No wildcards. 3. **Set approval gates for destructive operations.** Any operation that's hard to reverse — pushing to main, deleting resources, sending external communications — should require a human click in Slack before proceeding. 4. **Configure token budgets and alerts.** Set a daily spend limit. Alert when the agent hits 50% of budget. Kill the session if it hits 100%. 5. **Enable full session logging.** Every action, every result, timestamped and queryable. With these controls in place, YOLO mode isn't reckless. It's a fast, autonomous agent with a control plane watching everything it does. ## The Name "YOLO" is a joke, but the underlying concept is serious. You Only Live Once — act boldly, move fast. The irony is that YOLO mode is only genuinely safe when you've done the careful, deliberate work of building a control layer around it. Speed and safety aren't opposites when you architect for both. --- ### AI Agent Governance: Building Control Into Your Claude Agent Stack **URL:** https://sentrely.com/blog/ai-agent-governance **Published:** 2026-04-21 **Description:** 75% of enterprise AI projects fail at operationalization. Governance isn't the problem — bad governance is. Here's how to build it in from day one. The statistic worth taking seriously: 75% of enterprise AI projects that reach pilot stage fail to reach production. The most common reason isn't the model. It isn't the data. It's operationalization — the unglamorous work of making AI systems safe enough, auditable enough, and controllable enough to actually run in a real business context. Governance is what operationalization looks like for AI agents. ## A named category, not a nice-to-have AI agent governance is now its own analyst category, not a vague aspiration. Gartner's 2026 Hype Cycle for Agentic AI tracks it; Forrester projects a majority of the Fortune 100 will appoint a head of AI governance during 2026; and industry analysts now list "AI agent governance platforms" as a distinct software category — separate from MLOps and LLM observability — defined by exactly the capabilities below: an immutable, legal-grade log of every tool call, fine-grained per-action permissions, human-in-the-loop approval, and strict control over which tools an agent can invoke. The moment the category got a name, "we'll add governance later" stopped being a defensible plan. ## Why AI Governance Fails Most teams approach governance the wrong way. They build the agent, prove it works, then try to add governance on top. By that point, governance looks like friction — approval gates that slow things down, audit requirements that add engineering work, access controls that break capabilities the agent depended on. The result is governance theater: a compliance checkbox with no real control. Or worse, governance gets abandoned entirely because it's too hard to retrofit. The second failure mode is the wrong optimization target. Teams measure governance success by whether auditors are happy, not by whether they can actually control their agents. These are related but not the same thing. Real governance means you can answer "what is agent X doing right now?" and "what exactly happened in session Y?" in under 60 seconds. The third failure mode is committee paralysis. Governance becomes a committee that slows every decision to a crawl. This is governance as bureaucracy rather than governance as engineering. ## The Four Pillars Effective AI agent governance rests on four capabilities. They're not novel concepts — they map directly to what good operations looks like in any system. **Transparency: every action is visible.** An agent that does things you can't see is not a system you control — it's a liability you operate. Transparency means every tool call, every API request, every decision point is logged with enough context to understand what happened and why. Not sampled. Not summarized. Every one. This is more demanding than it sounds. In a busy multi-agent environment, you might have thousands of actions per hour. The question isn't whether you can capture them — it's whether you can query them usefully when something goes wrong. **Accountability: every action has an owner.** In a well-governed agent system, every action is attributable to a specific agent identity operating under a specific policy in a specific session. Not "an agent did this." `claude-deploy-01` in session `d97e2169` operating under policy `project-a-v2` did this, and here's the full transcript. This requires per-agent identity (no shared credentials) and session tracking. Both are easy to skip and painful to add later. **Control: you can change what agents do.** Governance isn't worth much if you can't act on what you observe. Control means you can modify an agent's policy, revoke its access, kill its session, or pause all agents in a project — and have those changes take effect immediately, not on the next deployment. Policy-as-code makes this tractable. A YAML file that defines what each agent can do, committed to version control, enforced by a gateway. Changing the policy changes the agent's behavior without redeployment. **Monitoring: the system tells you when something's wrong.** Humans shouldn't have to watch dashboards to catch agent problems. The system should surface anomalies proactively: a loop that looks stuck, a budget that's burning faster than expected, an agent making requests outside its normal pattern. This is the layer most teams skip. It's also the layer that would have caught most major agent incidents. ## Policy-as-Code: A Practical Example Here's what a real agent policy looks like: ```yaml project: billing-automation agents: claude-invoice-01: allow: - service: git actions: [push, pull] resources: ["acme/billing", "acme/invoices"] branches: ["feature/*", "fix/*"] - service: aws actions: ["s3:GetObject", "s3:PutObject"] resources: ["arn:aws:s3:::invoice-data/*"] - service: stripe actions: ["charges:read", "invoices:read"] require_approval: - service: git actions: [push] branches: ["main", "release/*"] - service: stripe actions: ["charges:create", "refunds:create"] budget: daily_tokens: 500000 alert_threshold: 0.8 ``` This policy is readable by non-engineers, reviewable in a PR, and enforced consistently regardless of what the agent decides to do. The agent can push to feature branches without approval. It cannot push to main without a human sign-off. It can read Stripe data but cannot create charges without approval. Change the policy file, and the agent's behavior changes immediately on the next request. No redeployment. ## Autonomy vs. Oversight: The Risk Tier Model Not every operation needs the same level of oversight. A good governance framework tiers operations by risk: **Tier 1 — Run freely.** Low blast radius, fully reversible. Reading files, querying databases (read-only), fetching APIs. These happen thousands of times per session and don't need human review. **Tier 2 — Log and alert.** Medium blast radius, reversible with effort. Writing files, pushing to non-protected branches, writing to development databases. These run autonomously but generate audit events that can be reviewed. **Tier 3 — Require approval.** High blast radius or hard to reverse. Production deployments, data deletion, sending external communications, large financial transactions. These pause and route to a human before proceeding. **Tier 4 — Never.** Outside policy regardless of context. Deleting production databases, modifying IAM policies, accessing data outside project scope. These are denied at the gateway. The mistake most teams make is applying the same tier to everything — either "approve everything" (kills velocity) or "approve nothing" (provides no protection). ## Common Mistakes **Shared credentials.** One API key for 10 agents means you can't attribute actions, you can't revoke one agent without breaking all of them, and your audit trail is meaningless. **After-the-fact governance.** Adding governance after the agent is built means fighting against dependencies. Build the policy layer first, then build the agent inside it. **Monitoring without alerting.** A dashboard nobody watches is not monitoring. Governance needs to surface problems to humans proactively, in the channels they already use. **Over-broad policies.** "Allow everything except deleting databases" is not a policy. It's an invitation for incidents you haven't imagined yet. Start with least privilege and expand as needed. ## Implementation Roadmap **Week 1: Inventory.** What agents are running? What credentials do they have? What can they actually do? Most teams are surprised by the answer. **Week 2: Policy definition.** Write policies for each agent based on what it actually needs (not what it has). This is the hardest week — expect pushback from teams used to broad access. **Week 3: Gateway deployment.** Route all agents through a control plane. Enforce policies. Start collecting audit data. **Week 4: Approval gates and alerting.** Configure approval requirements for high-risk operations. Set up budget alerts. Wire notifications to Slack. After week 4, you have a governed agent stack. Not perfect, but real. The difference between "we have AI agents" and "we run AI agents safely" is measured in these four weeks. ## Where to go next Governance is the pillar; these are the load-bearing parts: - [What is an AI agent control plane?](/blog/what-is-an-ai-agent-control-plane) — the enforcement layer that makes governance real rather than documented. - [Claude Code security: what the native permissions don't cover](/blog/claude-code-security) — why agent config alone fails open. - [Why static RBAC isn't enough for AI agents](/blog/rbac-for-ai-agents) — deny-by-default, scoped to action + resource. - [Approval gates](/blog/claude-code-approval-gates) — the human-in-the-loop tier. - [Immutable audit trails](/blog/claude-agents-audit-trail) — accountability and compliance evidence. ## FAQ **What is AI agent governance?** The practice (and now a named software category) of making autonomous AI agents safe to run in production: transparency (every action logged), accountability (every action attributable to an agent identity + policy + session), control (change policy, revoke access, kill sessions immediately), and proactive monitoring. It maps to capabilities like immutable audit, fine-grained permissions, and human-in-the-loop approval. **What is an AI agent governance platform?** Software that enforces those capabilities through a control plane the agent can't bypass — deny-by-default policy on every action, per-agent identity, an immutable audit trail, approval gates, and cost controls — rather than relying on the agent's own configuration. **How is AI agent governance different from LLM observability?** Observability tells you what happened (traces, tokens, latency). Governance also *prevents* it: it enforces policy before an action runs and can stop it. Analysts treat them as distinct categories for this reason. **When should we add governance?** Before the agent ships, not after. Retrofitting governance means fighting the dependencies the agent already built on broad access — the most common reason agent projects stall between pilot and production. --- ### What Is an AI Agent Control Plane? The Missing Layer in Your AI Stack **URL:** https://sentrely.com/blog/what-is-an-ai-agent-control-plane **Published:** 2026-04-20 **Description:** An AI agent control plane is the policy-enforcing gateway between your agents and everything they touch. Learn what it does, why routing Claude Code through a central gateway beats distributed API keys, and why every production team needs one. If you've spent time with Kubernetes, you know what a control plane is: the layer that watches everything, enforces policy, and makes sure the system does what you intend rather than what it wants. Worker nodes do the compute; the control plane keeps them from going rogue. AI agents need the same thing. And almost nobody has built it yet. ## The gap between demo and production Running a Claude agent in demo mode is easy: open a terminal, run `claude`, watch it write code, call APIs, push commits. Impressive. Feels safe. Then you run it in production — against real infrastructure, real data, real systems — and the gaps appear fast: - Which AWS actions is this agent allowed to take? - Who approved that git push to `main`? - How much has this agent spent in the last 24 hours? - When it hit a 403 on that S3 bucket, what did it do next? - Can you replay what happened in session `d97e2169`? Without a control plane, the answer to all of these is "we don't know." ## What a control plane actually does A control plane for AI agents sits between your agents and everything they can touch. Every tool call, API request, and git operation routes through it. It does five things: **1. Policy enforcement (RBAC).** Every agent operates under a [deny-by-default policy](/blog/rbac-for-ai-agents). `project-a` can push to `feature/*` but not `main`; call `s3:GetObject` on `data-lake/raw/*` but not `s3:DeleteObject`; invoke Bedrock but only with specific model IDs. Enforced on every request by interception — not by trusting the agent. **2. Full audit trail.** Every action logged: what the agent tried, whether it was allowed, the result, the duration. Not optional telemetry — the foundation of accountability. When something breaks, you answer "what exactly happened?" in under 60 seconds. **3. Human-in-the-loop approvals.** Some operations shouldn't run autonomously regardless of policy — pushing to prod, dropping a table, emailing 10,000 customers. The control plane [gates these](/blog/claude-code-approval-gates) and routes approval to a human via Slack, Telegram, or dashboard. **4. Cost controls.** Agents are prolific. A loop bug burns $200 before a human notices. Per-project token budgets, rate limits, and spend alerts are the difference between an incident and a surprise invoice. **5. Session and agent identity.** A shared API key across 12 agents makes your audit trail useless. The control plane assigns each agent a session identity and scopes permissions to it. ## Why a central gateway beats distributed API keys This is the architectural decision underneath all of it, and it's not just our opinion — **Anthropic's own Claude Code documentation recommends routing through an LLM gateway**: a centralized proxy that provides "single-point API key management, usage tracking, cost controls, and audit logging." The alternative — handing every developer and every agent its own API key and credentials — gives you N copies of the credential-leak problem and zero central visibility. Routing all Claude Code traffic through one control plane means: - API keys and cloud credentials live in **one** governed place, not scattered across laptops and CI runners - Every agent inherits policy, audit, and cost controls automatically — no per-agent setup - You can rotate keys, change policy, or kill a session **without touching the agents** This is why competing vendors (TrueFoundry, Kong, Portkey) and Anthropic itself all converge on the same pattern. The difference is what the gateway governs: most stop at routing and cost. A true control plane adds per-agent RBAC, human approvals, and an immutable audit trail. ## What happens without one Three real failure modes: - **The runaway loop.** An agent retrying a failed call with no circuit breaker — 73 calls in 19 minutes, all `429`, $4.20 burned, unnoticed for half an hour. With a control plane, the watchdog pinged Slack at call 10. - **The blast-radius incident.** An agent with broad AWS permissions told to "clean up old resources" interpreted it liberally — S3 objects, log groups, IAM roles, gone. No audit trail, no approval gate, no way to know what was deleted. - **The late-night deploy.** An agent with push access to all branches decided a refactor was ready and pushed to `main` overnight. The pipeline ran. The bug hit prod at 3 AM. All three are prevented by a control plane with a few lines of YAML policy and an approval gate. ## Control plane vs. raw API access Direct API access — give the agent credentials, let it call whatever — is fine for a single trusted developer building a private tool. It breaks down the moment you have: - Multiple agents in one environment - Any agent touching production - Any security or regulatory requirement - Any need to understand what your agents are actually doing The control plane isn't bureaucracy. It's the operating model that makes autonomous agents safe to run at meaningful scale. (See also: [Claude Code security — what the native permissions don't cover](/blog/claude-code-security).) ## The Sentrely approach Sentrely is a managed control plane for Claude Code and Codex agents. You get a dedicated endpoint at `you.sentrely.io`, define policies in YAML, and point your agents' gateway URL at it. Every agent call routes through the gateway, which enforces your policy, logs everything immutably, and routes approvals to your Slack or Telegram. The agent holds **zero credentials** — it asks the gateway, the gateway decides, the gateway acts. No infrastructure to run. No Postgres to configure. No gateway to patch. You get the speed of YOLO mode with the safety of a production control system. If you're running Claude agents against anything that matters, you need a control plane. The only question is whether you build it or use one that's already built. ## FAQ **What is an AI agent control plane?** A policy-enforcing layer that sits between your AI agents and the systems they access. Every tool call, API request, and command routes through it; it checks each against a deny-by-default policy, logs it immutably, gates risky operations for human approval, and enforces cost limits — so the agent operates within bounds you control rather than the full reach of its credentials. **What's the difference between an AI gateway and a control plane?** An AI/LLM gateway centralizes API traffic — key management, routing, cost, basic logging (Anthropic recommends one for Claude Code). A control plane adds the governance layer on top: per-agent RBAC, human-in-the-loop approvals, immutable per-action audit, and a kill switch. Every control plane is a gateway; not every gateway is a control plane. **Should I route Claude Code through a gateway?** Yes — Anthropic's own docs recommend it for centralized key management, cost control, and audit logging. Routing all traffic through one gateway also lets you enforce policy and kill sessions without reconfiguring each agent. **Do I have to build my own control plane?** No. You can assemble one from IAM, a proxy, CloudTrail, and approval tooling — or use a managed control plane (like Sentrely) that provides policy enforcement, audit, approvals, and cost controls out of the box with no infrastructure to run. --- ## Playbooks ### Compliance Readiness for Claude Agents **URL:** https://sentrely.com/playbooks/compliance-readiness **Subtitle:** How to make your Claude agent operations audit-ready for SOC 2, HIPAA, GDPR, and beyond ## Chapter 1: Which Frameworks Apply? **SOC 2 Type II** — You sell software to businesses and customers ask about your security practices. If you've ever received a security questionnaire from a prospect, you need SOC 2. It's the baseline for B2B trust. **HIPAA** — You handle Protected Health Information in any form. If agents process medical records, patient data, insurance claims, or health-related data, HIPAA applies regardless of company size. **GDPR** — You process personal data of EU residents. Applies regardless of where your company is located. **PCI-DSS** — You process, store, or transmit credit card data. Most startups need SOC 2 first, GDPR if they have EU customers, HIPAA only if in healthcare. ## Chapter 2: SOC 2 Type II SOC 2 evaluates five Trust Service Criteria. Sentrely satisfies each: **Security (required):** Per-agent RBAC, session tokens with automatic expiration, deny-by-default, approval gates for privileged operations. **Processing Integrity:** Every agent action logged with timestamp, identity, resource, and outcome. Data pipelines include validation gates. **Confidentiality:** Per-project isolation prevents cross-project access. Scoped credentials mean agents only access what they need. **Evidence for auditors:** - Policy files showing per-agent permissions - Audit logs showing denied events (proves enforcement, not just declaration) - Approval logs showing human-in-the-loop for privileged operations ## Chapter 3: HIPAA for AI Agents HIPAA adds specific technical safeguards: **Access Control (164.312(a))** — Unique user identification and automatic logoff. → Sentrely: Per-agent identity with session tokens that expire automatically. **Audit Controls (164.312(b))** — Record and examine activity in PHI-containing systems. → Sentrely: Complete audit trail of every data access, automatically generated. **Transmission Security (164.312(e))** — Protect PHI during transmission. → Sentrely: All communication encrypted (TLS). Gateway-to-agent traffic is HTTPS only. **VPC Deployment:** For HIPAA, deploy Sentrely within your VPC. Enterprise tier supports VPC deployment where no PHI leaves your environment. **BAA Requirement:** If using Sentrely's managed cloud with PHI, a Business Associate Agreement is required. Contact sales. ## Chapter 4: GDPR Data Residency and Erasure **Data Residency (Articles 44-49):** EU resident data must stay in the EU or adequately protected countries. ```yaml data_residency: region: eu-central-1 audit_storage: s3://eu-audit-bucket/ retention: personal_data_logs: 365d anonymized_metrics: 2555d ``` **Right to Erasure (Article 17):** Audit logs containing personal data must be deletable on request. Configure retention policies that automatically purge old records. Session data that processed personal data must be traceable so you can identify and delete relevant logs. ## Chapter 5: The Universal Audit Trail One well-designed audit trail satisfies SOC 2, HIPAA, and GDPR simultaneously: | Field | SOC 2 | HIPAA | GDPR | |-------|-------|-------|------| | Timestamp | ✓ | ✓ | ✓ | | Agent identity | ✓ | ✓ | ✓ | | Action type | ✓ | ✓ | ✓ | | Resource accessed | ✓ | ✓ | ✓ | | Outcome (allowed/denied) | ✓ | ✓ | ✓ | | Policy applied | ✓ | — | ✓ | | Data category (PHI/PII) | — | ✓ | ✓ | | Session ID | ✓ | ✓ | — | Configure Sentrely: ```yaml audit: fields: [timestamp, agent, action, resource, outcome, policy, data_category, session_id] export: destination: s3://audit-bucket/ format: json encryption: AES-256 frequency: hourly ``` ## Chapter 6: Preparing for an Audit Auditors want to see three things: 1. **Policies** — Human-readable YAML showing what each agent can and cannot do 2. **Enforcement evidence** — Denied requests in the audit log (proves policies are enforced, not just declared) 3. **Incident response** — Walk through what happens when something goes wrong The most powerful moment in an audit: when the auditor asks "what would happen if an agent tried to access data outside its scope?" and you can open the audit log and show them a previous denied request with timestamp, agent identity, and the policy that blocked it. ## Chapter 7: Ongoing Compliance Operations **Monthly:** Review agent policies for accuracy. Check audit log completeness. Review approval gate effectiveness. **Quarterly:** Run a mock audit. Update data retention policies. Verify log exports are working. **Annually:** Full policy review, remove unused agents and stale permissions. Assess whether new frameworks apply. **On policy change:** Document the reason, log the approver, verify the new policy works via testing. The goal: compliance evidence generated automatically as a byproduct of normal operations — not assembled manually before each audit. --- ### Claude API Cost Optimization Playbook **URL:** https://sentrely.com/playbooks/cost-optimization **Subtitle:** Token budgets, rate limits, and architecture patterns that reduce Claude API spend by 40-60% ## Chapter 1: Where Your Tokens Actually Go Most teams are surprised by the answer. The three biggest token sinks: **Context length** — Every request sends the full conversation context. A session starting at 2,000 tokens might send 50,000 tokens per request by step 50. You pay for the full context on every single call. **Retry loops** — When an agent hits an error, it retries. Without a circuit breaker, it retries indefinitely. 100 failed calls cost almost as much as 100 successful ones, and produce nothing. **Parallel sessions** — Five agents at 100k tokens each is 500k tokens. If they're all running long sessions simultaneously, actual costs are higher than the sum. Use the Gateway's Cost Analytics to see your actual distribution: what percentage goes to context vs. new completions? Where are the cost spikes? ## Chapter 2: Per-Project Budgets and Per-Session Caps **Per-session caps** are your primary control: ```yaml budget: max_tokens_per_session: 100000 ``` How to set: run your agent 10 times on typical tasks, measure usage, set cap at 2x the average. | Agent Type | Typical Session | Recommended Cap | |------------|-----------------|-----------------| | Code review | 20k-50k | 100k | | CI/CD pipeline | 30k-80k | 200k | | Research | 50k-150k | 300k | | Support ticket | 10k-30k | 50k | **Per-project monthly budgets** prevent aggregate cost creep: ```yaml project_budget: monthly_limit: $500 alert_thresholds: [50%, 75%, 90%] on_exceed: pause_and_alert ``` The 50% alert is the most useful — it fires with time to investigate before hitting the limit. ## Chapter 3: Rate Limiting to Catch Loops Rate limiting is your circuit breaker for retry loops: ```yaml rate_limits: max_requests_per_minute: 30 burst_allowance: 10 on_exceed: throttle_and_alert ``` Most legitimate agent work averages 5-15 requests/minute. Thirty is generous for bursts but catches loops firing every second. The retry loop detector catches the most expensive failure mode: ```yaml alerts: - type: retry_loop_detected threshold: 10_identical_requests channel: slack:#agent-alerts action: throttle ``` When 10 identical requests appear in sequence, something's wrong. Gateway throttles + alerts, giving you time to investigate. ## Chapter 4: Context Optimization **Shorter system prompts** — Sent on every request. A 5,000-token system prompt on 100 requests = 500,000 tokens in system prompt alone. Trim to essentials. **Task chunking** — Instead of one long session reviewing 50 files, run five sessions of 10 files each. Each starts with fresh context: | Approach | Tokens | Cost | |----------|--------|------| | One session, 50 files | ~2,000,000 | $6.00 | | Five sessions, 10 files | ~500,000 | $1.50 | Chunking uses 75% fewer tokens because each session starts fresh rather than accumulating all previous context. ## Chapter 5: Monitoring and Alerting ```yaml alerts: - type: budget_threshold level: project thresholds: [50%, 75%, 90%] channel: slack:#cost-alerts - type: session_cost_spike threshold: 3x_average channel: slack:#cost-alerts - type: context_growth_anomaly threshold: 5x_initial_context channel: slack:#agent-alerts ``` ## Chapter 6: Cost Attribution in Multi-Agent Setups The Gateway tracks cost at four levels: per-request, per-session, per-agent, per-project. Use this to find your optimization targets. Common findings: - Research agent costs 5x more than others (processes large documents) - CI/CD agent has occasional 10x spikes from large diffs - Support agent is cheap per session but runs 200/day = highest aggregate ## Chapter 7: Expected Results | Optimization | Typical Savings | |-------------|-----------------| | Per-session caps (catching runaways) | 15-25% | | Rate limiting (catching retry loops) | 10-20% | | Task chunking (shorter contexts) | 15-25% | | System prompt optimization | 5-10% | | Removing unused agents | 5-15% | **Total: 40-60% cost reduction** for teams implementing all optimizations. Start with budgets and rate limits — they catch the worst problems immediately and require minimal configuration. --- ### Enterprise Claude Agent Deployment **URL:** https://sentrely.com/playbooks/enterprise-deployment **Subtitle:** VPC isolation, SSO, cross-account IAM, and scaling from 10 to 100+ agents ## Chapter 1: Enterprise Requirements Enterprise deployments have requirements that go beyond standard managed cloud: **VPC isolation** — The Gateway runs within the customer's AWS VPC. Data doesn't leave the customer network. Non-negotiable for regulated industries and most large enterprises. **SSO/SAML** — Engineers authenticate with the company's IdP (Okta, Azure AD, Google Workspace). When someone leaves, access is revoked automatically. **Cross-account IAM** — Large enterprises operate across multiple AWS accounts. Agents need scoped access across accounts without permanently stored credentials. **Dedicated SLA** — Production agent operations need an availability commitment and a direct support channel. **Change management** — New agents and policy changes go through security, compliance, and engineering leadership review. ## Chapter 2: VPC Architecture Sentrely Enterprise deploys as containerized services within the customer VPC: ``` Customer VPC ├── Sentrely (ECS/EKS) ├── Agent Pool (containers) ├── Audit Store (S3 + DynamoDB — append-only) ├── Policy Store (DynamoDB) └── Dashboard (CloudFront/ALB — SSO authenticated) ``` All components run within the VPC. The Dashboard is accessible through an internal ALB, authenticated via SSO. No data leaves the customer environment. ## Chapter 3: Cross-Account IAM Setup **Step 1: Gateway role** in the Gateway account: ```json { "Statement": [{ "Effect": "Allow", "Action": "sts:AssumeRole", "Resource": [ "arn:aws:iam::DEVACCOUNT:role/YoloGateway-Dev", "arn:aws:iam::STAGINGACCOUNT:role/YoloGateway-Staging", "arn:aws:iam::PRODACCOUNT:role/YoloGateway-Prod" ] }] } ``` **Step 2: Target account roles** in each account, trusting the Gateway role with an external ID (prevents confused deputy attacks). **Step 3: Per-agent role mapping:** ```yaml agent: deployer cross_account: - account: "STAGINGACCOUNT" role: YoloGateway-Staging allowed_actions: [ecs:UpdateService, ecr:PutImage] - account: "PRODACCOUNT" role: YoloGateway-Prod allowed_actions: [ecs:UpdateService] requires_approval: true ``` The Gateway assumes the appropriate role per request and vends temporary credentials to the agent. Credentials expire with the session — no long-lived keys anywhere. ## Chapter 4: SSO/SAML Integration Supported providers: Okta, Azure AD, Google Workspace, OneLogin, any SAML 2.0 IdP. ```yaml auth: type: saml idp_metadata_url: https://yourcompany.okta.com/app/metadata attribute_mapping: email: claims/emailaddress groups: claims/Group role_mapping: admin: "Engineering-Leads" operator: "Engineering" viewer: "Engineering-All" ``` | SSO Group | Gateway Role | Permissions | |-----------|-------------|-------------| | Engineering-Leads | Admin | Create projects, manage policies, approve requests | | Engineering | Operator | View projects, manage own agents, approve within scope | | Engineering-All | Viewer | View dashboards and audit logs | ## Chapter 5: Scaling from 10 to 100+ Agents **10-25 Agents (Team Level)** - Establish naming conventions - One engineer per team as "agent owner" - Set up per-team Slack channels for notifications **25-50 Agents (Department Level)** - Policy templates so teams don't reinvent common patterns - Cost allocation tags for finance reporting - Approval delegation model (team leads within their scope, security for cross-team) - Self-service onboarding via form that generates policies from templates **50-100+ Agents (Organization Level)** - Dedicated platform team owns Gateway infrastructure - Policies stored in version control, reviewed via PR, applied via CI/CD - Automated compliance reporting (monthly SOC 2 evidence packages from audit logs) - Per-team cost chargeback in finance dashboards - Agent incident response playbook ## Chapter 6: Change Management The biggest challenge isn't technical — it's organizational. Getting teams to adopt governance feels like adding friction. **Start with value, not rules:** "Governance lets you do MORE with agents. Without it, security blocks expansion. With it, you can deploy agents for production workloads." **Phased rollout:** 1. **Observe** — Deploy in monitoring mode. Log everything, deny nothing. Show teams what their agents are doing. 2. **Recommend** — Enable enforcement with a 7-day warning period. Show what would be denied. 3. **Enforce** — Turn on enforcement. Policies are tuned, teams understand the model. 4. **Expand** — With governance in place, approve new high-stakes use cases. **The Governance Champion:** One senior engineer per team who writes policies, reviews changes, trains teammates, and escalates issues. ## Chapter 7: Measuring Success **Operational:** Agent uptime (target: 95%+), time to deploy a new agent (target: under 4 hours with templates), incident rate (target: under 0.1 per agent-month). **Security:** Policy violation rate (should decrease as policies mature), mean time to detect violations (target: under 5 minutes). **Business:** Developer time saved per agent per week, time to SOC 2 evidence generation (target: automated, under 1 hour). **Adoption:** Teams using governed agents, agents in production with growth trend, percentage of operations going through the Gateway (target: 100%). Report these monthly. The narrative: governed agents are safer, cheaper, and more productive than ungoverned ones — and the data proves it. --- ### Your First Claude Agent in Production **URL:** https://sentrely.com/playbooks/first-agent-in-production **Subtitle:** From zero to a controlled, auditable Claude agent in one afternoon ## Chapter 1: Pick the Right First Use Case Not every task suits a first agent deployment. Choose something that is **bounded** (clear, limited access), **measurable** (you know when it worked), **non-critical** (failure doesn't break production), and **repetitive** (happens often enough to justify setup). Good first choices: code review on one repo, daily ticket summary, documentation generation. Avoid: production deployments, customer-facing interactions, multi-system orchestration. ## Chapter 2: Set Up Sentrely Log in to the Sentrely dashboard and create a project. Note your `GATEWAY_URL` — every agent request routes through it. ```bash export GATEWAY_URL=https://gw.yologateway.io/projects/my-webapp ``` ## Chapter 3: Write Your First Policy For a code review agent: ```yaml project: my-webapp agent: code-reviewer policies: - git:read on repos/my-webapp - git:comment on repos/my-webapp/pull-requests/* # Nothing else. No push, merge, deploy, or other repos. budget: max_tokens_per_session: 100000 max_sessions_per_day: 50 ``` Apply via the dashboard. The policy is active immediately — any request outside these permissions is denied and logged. ## Chapter 4: Run Your First Session Prepare a session token: ```bash curl -X POST $GATEWAY_URL/sessions/prepare \ -H "Authorization: Bearer $ADMIN_TOKEN" \ -d '{"project": "my-webapp", "agent": "code-reviewer"}' ``` Launch Claude with the session token: ```bash docker run --rm \ -e GATEWAY_URL=$GATEWAY_URL \ -e SESSION_TOKEN=$SESSION_TOKEN \ -e ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \ your-agent-image:latest ``` ## Chapter 5: Review the Audit Trail In the Gateway dashboard → **Audit Log**, filter by project and agent. You'll see every request: - `git:read` on `repos/my-webapp/src/main.ts` — Allowed ✓ - `git:comment` on `pull-requests/42` — Allowed ✓ - `git:push` on `branches/main` — **Denied** ✗ (policy violation) That denied push means the Gateway worked. The agent tried to exceed its bounds and was stopped. ## Chapter 6: Add a Slack Approval Gate Update your policy to gate security findings: ```yaml policies: - git:read on repos/my-webapp - git:comment on repos/my-webapp/pull-requests/* condition: comment_tag: security requires_approval: true approval_channel: slack:#security-reviews ``` Security-tagged comments now pause and route to Slack. A team member approves or denies with one click. The decision is logged. ## Chapter 7: What's Next You have a governed Claude agent. From here: 1. **Expand policies** — add repos, new capabilities 2. **Add more agents** — each with its own identity and policy 3. **Set up cost alerts** — catch unexpected token usage 4. **Browse templates** — pre-built configs for common use cases 5. **Read the security playbook** — advanced RBAC and compliance Governance doesn't slow you down. It lets you run more agents with more confidence. --- ### Multi-Agent Architecture Playbook **URL:** https://sentrely.com/playbooks/multi-agent-setup **Subtitle:** Patterns, policies, and operational practices for running multiple Claude agents safely ## Chapter 1: When to Go Multi-Agent **Go multi-agent when you need:** - **Parallelism** — independent tasks that should run simultaneously - **Specialization** — different tasks need different expertise and different permissions - **Isolation** — failure in one task shouldn't affect another - **Audit clarity** — clean, attributable trails per agent vs. tangled monolith **Stick with one agent when** you only have one use case. Multi-agent adds coordination complexity — only split when the benefit outweighs it. ## Chapter 2: Architecture Patterns **Pipeline (sequential):** Agent A → Agent B → Agent C. Each agent's output becomes the next agent's input. Use when tasks have natural dependencies. ``` Code Review Agent → Deploy Agent → Notification Agent ``` **Fan-out (parallel):** One coordinator spawns multiple workers in parallel, collects results. ``` PR Trigger → [Security Scanner + Code Reviewer + Docs Checker] → Summary ``` **Hierarchical:** A supervisor agent delegates to specialists based on context. ``` Research Coordinator → [Data Gatherer + Analysis Agent + Report Writer] ``` ## Chapter 3: Per-Agent Isolation Every agent needs its own identity and policy. Never share identities between agents. **Same project, different agents:** ```yaml # Project: acme-webapp agent: code-reviewer policies: - git:read on repos/acme-webapp - git:comment on repos/acme-webapp/pull-requests/* --- agent: deployer policies: - git:push on repos/acme-webapp/branches/feature/* - aws:ecs:UpdateService on arn:aws:ecs:*:*:service/staging-acme-* ``` **Separate projects** when you need hard isolation — agents in different projects cannot see each other's resources, sessions, or audit trails. ## Chapter 4: Agent-to-Agent Messaging Sentrely provides A2A messaging so agents coordinate without manual intervention: ```bash # Agent A sends a message to Agent B curl -X POST $GATEWAY_URL/a2a/send \ -H "Authorization: Bearer $SESSION_TOKEN" \ -d '{ "to": "deployer", "type": "pipeline_stage_complete", "payload": {"stage": "code_review", "result": "passed", "pr": 142} }' ``` A2A messages are logged in the audit trail. Control which agents can message which via policy: ```yaml agent: code-reviewer policies: - a2a:send to deployer - a2a:send to notification-bot # Cannot message: support-bot, security-auditor ``` ## Chapter 5: Fleet Monitoring With multiple agents you need a fleet-level view. Sentrely dashboard shows: - **Agent Status** — Active, idle, or failed sessions in real-time - **Resource Usage** — Token consumption per agent with trend lines - **Audit Timeline** — Unified view of all agent actions, filterable - **Alert Feed** — Denied requests, budget warnings, approval requests Set up a daily digest: total sessions, total tokens (with cost), approval outcomes, denied requests, errors. ## Chapter 6: Cost at Scale Set budgets at three levels: ```yaml budget: per_session: 100000 tokens # Catches runaway sessions per_agent_per_day: 500000 tokens # Catches over-active agents per_project_per_month: $500 # Catches aggregate creep ``` Cost attribution identifies which agents cost most. Often, one or two agents account for 80% of total spend — focus optimization there. ## Chapter 7: Common Mistakes **Starting with too many agents** — Begin with one, get governance right, add a second, stabilize, then expand. **Shared credentials** — Every agent needs its own identity. Shared credentials make audit trails useless. **No communication boundaries** — Define the communication graph explicitly. Not every agent should be able to message every other. **No kill switch** — You need to stop any agent immediately. Test the session termination function before you need it. **One big policy** — Each agent gets its own policy. When you read a policy, you should understand exactly what that one agent can do. --- ### Securing Your Claude Agent Stack **URL:** https://sentrely.com/playbooks/securing-your-claude-stack **Subtitle:** From shared credentials and no audit trail to a hardened, compliant agent fleet ## Chapter 1: Audit Your Current State Answer these honestly: - Are your agents sharing a single API key or AWS credential set? - Can you tell which agent used which credential and when? - Does every agent have the same level of access? - Can you reconstruct what each agent did last week? - Are any operations gated on human approval? If most answers are "no," you're in the common starting position — and one incident away from a bad day. ## Chapter 2: Implement Least-Privilege RBAC **Step 1: Map agents to their actual needs** | Agent | Needs | Does NOT Need | |-------|-------|---------------| | code-reviewer | Repos (read), PR comments (write) | Push, merge, deploy, databases | | ci-runner | Repos (read), ECR (push), ECS staging | Production, IAM, databases | | support-bot | Tickets (read/write), CRM (read) | Code repos, infrastructure | **Step 2: Write per-agent policies** ```yaml # code-reviewer — read and comment only policies: - git:read on repos/* - git:comment on repos/*/pull-requests/* # ci-runner — staging free, production gated policies: - git:read on repos/* - git:push on repos/*/branches/feature/* - aws:ecs:UpdateService on arn:aws:ecs:*:*:service/staging-* - aws:ecs:UpdateService on arn:aws:ecs:*:*:service/prod-* requires_approval: true ``` **Step 3: Test** — have each agent attempt an action outside its permissions. Verify it's denied and logged. ## Chapter 3: Set Up Audit Logging Sentrely logs everything automatically. Make it useful: **Configure retention** based on your requirements — SOC 2 needs 1 year, HIPAA needs 6 years. **Export to your SIEM:** ```yaml audit: export: destination: s3://your-audit-bucket/ format: json frequency: hourly ``` **Set up alerts** for: - Permission denied events (agent trying to exceed bounds) - 10x normal request volume (possible runaway loop) - Off-hours activity - New resource access patterns ## Chapter 4: Configure Approval Gates **Always gate:** production deploys, database migrations, customer data modification, financial transactions, IAM changes, secret rotation. **Usually safe to automate:** code review comments, test execution, staging deploys, internal notifications, documentation generation. ```yaml - aws:ecs:UpdateService on arn:aws:ecs:*:*:service/prod-* requires_approval: true approval_channel: slack:#deploy-approvals approval_timeout: 60m ``` ## Chapter 5: Set Token Budgets and Alerts ```yaml budget: max_tokens_per_session: 100000 # ~$0.30-1.50 depending on model max_sessions_per_day: 50 alert_at: 80% project_budget: monthly_limit: $500 alert_thresholds: [50%, 75%, 90%] on_exceed: pause_and_alert ``` Rate limiting catches runaway loops before they drain budget: ```yaml rate_limits: max_requests_per_minute: 30 on_exceed: throttle_and_alert ``` ## Chapter 6: Test Your Policies Run these prompts against each agent to verify controls: 1. Ask for an action outside its policy → should be denied 2. Ask for an action requiring approval → approval request should appear in Slack 3. Run a session approaching the token limit → alert should fire 4. Ask to access a different project's resources → should be denied Document results. Run monthly or after any policy change. ## Chapter 7: Ongoing Security Hygiene **Weekly:** Review denied requests and unusual patterns. **Monthly:** Audit agent permissions, remove access no longer needed. **Quarterly:** Rotate credentials, review approval gate effectiveness. **On incident:** Use the audit trail to trace the issue, identify the policy gap, update controls. ---