Alternatives · Trust & Compliance · v2026.Q3 · verified 2026-08-06

Promptfoo alternatives, 2026.Q3: every real option, ranked

The short answer

Edition v2026.Q3 · pricing and status verified 2026-08-06.

sandbox-runtime is the strongest Promptfoo alternative for most — putting a hard filesystem and network boundary around any agent process — coding agents especially — without containers or a cloud bill. Then: NVIDIA NeMo Guardrails · Amazon Bedrock Guardrails · Lakera Guard. Below, all 26 real options in the guardrails category, with pricing and honest watch-outs — plus the 8 "alternatives" other lists still recommend that are dead, renamed, or sunsetting.

Why people look past Promptfoo at all: It tests; it does not enforce. A green promptfoo run is not a runtime control, and nothing here blocks a live request. The OpenAI acquisition also puts the category's most-used neutral scanner inside a frontier lab — the OSS pledge is a pledge, not a governance structure.

The top Promptfoo alternatives, ranked

  1. 2sandbox-runtimeAnthropic (Apache-2.0, beta research preview)

    Free · Apache-2.0 · no service, no account

    Best for Putting a hard filesystem and network boundary around any agent process — coding agents especially — without containers or a cloud bill.

    Watch Explicitly a "beta research preview": Windows support is alpha, and the API is not stable. It does nothing about content — no injection detection, no PII, no topic control. And a sandbox with an over-broad network allowlist is theatre; the config is the product.

    4.6k stars, Apache-2.0, v0.0.64 released Jul 7, 2026 [src] · 84% reduction in permission prompts after OS-level sandboxing replaced approval-gating in Claude Code [src]
  2. 3NVIDIA NeMo GuardrailsNVIDIA (Apache-2.0)

    Free · Apache-2.0 · self-hosted; NemoGuard NIM microservices require NVIDIA AI Enterprise licensing (not publicly priced)

    Best for Teams that want programmable, self-hosted rails — including rails on tool calls — rather than a hosted moderation endpoint.

    Watch Colang is a language you have to learn, and the config surface sprawls fast. Every rail is another LLM call, so latency and token cost stack; teams routinely ship with rails disabled in the hot path. The best detectors are NVIDIA NIMs behind enterprise licensing, so the free tier is the scaffolding, not the models.

    6.9k stars, 802 forks, 3,765 commits, Apache-2.0, v0.23.0 (Aug 2026) [src] · Five rail types including execution rails that validate tool calls and external interactions [src]
  3. 4Amazon Bedrock GuardrailsAWS

    Usage-based, no subscription: content filters & denied topics $0.15/1k text units · sensitive-info filters $0.10 · contextual grounding $0.10 · Automated Reasoning $0.17 per policy · word filters and regex free · standalone InvokeGuardrailChecks prompt-attack $0.08/1k (1 text unit = 1,000 characters)

    Best for Shipping a managed filter today without building one — including in front of models you don't host on AWS.

    Watch The headline numbers — "blocks up to 88% of harmful content", "99% accuracy" on Automated Reasoning — are AWS's own, with no independent benchmark to check them against. Per-1k-character billing gets expensive on long agent contexts, every enabled policy is a separate charge and a separate round trip, and you're inside an AWS account. It filters text, not actions: it will not stop an agent from calling the wrong tool.

    ApplyGuardrail works with any foundation model including self-hosted and third-party (OpenAI, Google Gemini) — vendor page, Aug 2026 [src] · Exact list pricing: $0.15/1k text units content filters; $0.17/1k per Automated Reasoning policy; word filters free (AWS pricing page, Aug 2026) [src]
  4. 5Lakera GuardLakera → Check Point Software (announced Sep 16, 2025)

    Not published — enterprise via Check Point sales; SaaS and self-hosted deployment options

    Best for Startups selling into enterprises that want a named vendor, an SLA and a real adversarial dataset behind the injection filter.

    Watch No public pricing at all — the least startup-friendly entry on this list, and you're now buying from a large network-security vendor whose roadmap answers to CloudGuard, not to you. All performance figures are first-party. Detection percentages on prompt injection are also structurally soft: there is no accepted public benchmark, so >98% is a claim about Lakera's own test set.

    Gandalf network: 80M+ adversarial patterns; >98% detection, sub-50ms latency, <0.5% false positives (Check Point release, Sep 16, 2025) [src] · Vendor site (Aug 2026): 1M+ hackers, 100+ languages, sub-50ms latency, 0.01% production false-positive rate [src]

Every other live option in guardrails

All active guardrails tools beyond the top five — verified 2026-08-06.
ToolMakerWhat it isEntry
Google Model ArmorGoogle CloudRuntime screening for prompts, responses and agent interactions — injection/jailbreak, indirect injection, PII, malicious URLs, malware in files; inline for Gemini Enterprise Agent Platform and LangChain2M tokens/mo free, then $0.10/M
Meta LlamaFirewallMeta (PurpleLlama)The agent-native open guardrail: PromptGuard 2 (injection classifier), AlignmentCheck (audits agent chain-of-thought for goal hijacking), CodeShield (static analysis over 8 languages). Research-grade cadencefree (open weights)
Llama Guard 4 / Prompt Guard 2MetaOpen-weight input/output safety classifiers — the default self-hosted moderation models when you can't send text to a vendorfree (open weights)
Guardrails AIGuardrails AI, Inc.The original validator framework (7.1k stars, Apache-2.0, v0.10.2 Jun 4 2026) plus Guardrails Hub; company has broadened into an "AI reliability platform" around Snowglobe simulationfree OSS + Hub
OpenAI GuardrailsOpenAIMIT-licensed drop-in client wrapper (Python + TS): moderation, jailbreak, PII, URL allowlist, NSFW, off-topic, hallucination-vs-vector-store. Still labelled preview, ~223 stars — thin next to NeMofree (pays OpenAI API costs)
OpenAI Moderation APIOpenAIFree harm classifier — the zero-effort baseline, and still where a lot of teams stop; no injection or agent-action coveragefree
Cisco AI DefenseCisco (built on Robust Intelligence)Model validation + runtime guardrails + AI-aware SASE; expanded for "the agentic era" in Feb–Mar 2026. Ships an open mcp-scannerenterprise quote
ZenityZenityAgent-action control: reads agent intent and allows/modifies/blocks before execution across Copilot, ChatGPT Enterprise, Gemini, Claude, Bedrock and Vertex. $125M Series C led by Norwest, Aug 3 2026; 230+ staff, revenue tripled two years runningenterprise quote
HiddenLayerHiddenLayerModel scanning, AI detection & response; one of the oldest pure-play AI security vendors ($50M Series A, 2023) and now a rare uncaptured independententerprise quote
WitnessAIWitnessAIEmployee-facing AI usage policy and guardrails — governs which AI your staff use, adjacent to but not the same as bounding your own agentsenterprise quote
Pillar SecurityPillar SecurityEnd-to-end AI app security with runtime guardrails; $9M seed (Apr 2025), known for the "Rules File Backdoor" coding-agent researchenterprise quote
NeuralTrust TrustGateNeuralTrustOpen-source Go AI gateway with inline guardrails — the most-cited migration path for teams stranded by LLM Guard's archivalfree OSS + paid cloud
garakNVIDIAOSS LLM vulnerability scanner (~2.8k stars) — probe-based, CLI-first, complements rather than competes with promptfoofree
PyRITMicrosoft AI Red TeamPython risk-identification toolkit used by Microsoft's own red team; the most extensible of the OSS attack frameworks and the least turnkeyfree
DeepTeamConfident AIOSS red-teaming framework from the DeepEval team; strong OWASP-mapped attack coverage, tied to the DeepEval ecosystemfree
GiskardGiskardOSS testing + LLM scan with an EU-AI-Act framing; publishes the clearest independent OWASP Agentic Top 10 walkthroughsfree OSS + Hub
MindgardMindgardAutomated AI red-teaming as a service, Lancaster University spinout; continuous testing rather than one-off pentestenterprise quote
Gray Swan AIGray SwanPublic jailbreak arenas and paid red-team events used by frontier labs — the closest thing to a neutral adversarial benchmark suppliercontest + enterprise
E2BE2BFirecracker-microVM sandboxes for agent code execution — the hosted answer when sandbox-runtime's local model isn't enoughfree $100 credit · Pro $150/mo + $0.000014/vCPU-s
DaytonaDaytonaSub-second agent sandboxes, the main E2B alternative; competes on cold-start and per-second pricefree tier + usage
Docker MCP Gateway / RunlayerDocker · RunlayerMCP gateways that centralise auth, routing and allowlists for tool servers; Runlayer raised $11M in the thin MCP-security funding wave (~$40M total across four startups)free / usage
Auth0 for AI Agents / Permit.ioOkta · Permit.ioScoped, delegated credentials and fine-grained authorization for agents — the permission half of "scope it like a new hire", covered lightly by every guardrail framework herefree tier + usage

The "Promptfoo alternatives" to avoid — no longer what they were

Listicles still recommend these. As of 2026-08-06, they are not what the listicles think.

Former guardrails options — status verified 2026-08-06.
ToolStatusWhat happened
Microsoft Foundry guardrailsrenamedPrompt Shields, spotlighting, PII, groundedness plus agent-era previews: task adherence, tool-call and tool-response scanning, network egress controls. Azure AI Foundry renamed Microsoft Foundry in 2026
Prisma AIRSacquiredWhere Protect AI landed after the ~$500M acquisition completed Jul 22, 2025 — model scanning, runtime AI security, agent posture. Enterprise-only, no public pricing
Snyk Agent Scan / MCP-ScanacquiredThe team that publicly demonstrated MCP tool poisoning in Apr 2025, now Snyk Labs; scans MCP servers and agent skills for poisoned tool descriptions
Prompt SecurityacquiredGenAI prompt/response inspection folded into SentinelOne's platform two years after founding — an early sign the standalone LLM-firewall category wouldn't stay standalone
Aim SecurityacquiredAI security posture and runtime controls absorbed into a SASE platform — the same consolidation pattern as Lakera and Prompt Security
LLM GuarddeadThe category's most-forked OSS scanner suite — archived Jul 9, 2026, models on Hugging Face abandoned, final release v0.3.16 (May 2025). Migration path is Prisma AIRS; OSS successors are TrustGate and NeMo
RebuffdeadEarly prompt-injection detector (1.5k stars) — archived May 16, 2025, before the acquisition even closed
Robust IntelligencerenamedAcquired 2024; brand retired into Cisco AI Defense — the first of the AI-guardrail startups to be absorbed

How to choose

If your agent can run a shell, write files or reach the network
Sandbox first, filter second. Adversa's GuardFall work (Jun 2026) bypassed pattern-based command guards in 10 of 11 open-source coding agents; sandbox-runtime or an E2B-class VM sandbox is the boundary that doesn't depend on reading intent correctly.
If you have no guardrails and one afternoon
Run Promptfoo's red-team against your own agent and fix what it finds. It's free, self-hosted, and the exercise tells you which of the other four picks you actually need — instead of buying a filter for a hole you don't have.
If you want a managed filter and you're not on AWS
Bedrock's ApplyGuardrail API or Google Model Armor both work in front of third-party models. Model Armor is the cheaper start — 2M tokens/month free, then $0.10 per million — and covers indirect prompt injection and malicious URLs; Bedrock wins if you need denied topics or formally verified policy.
If your agent consumes MCP servers or third-party tools
Treat the tool registry as untrusted input: pin and scan servers (Snyk agent-scan, ex-Invariant MCP-Scan), gate them behind an MCP gateway, and put an allowlist on egress. GitGuardian found 24,008 secrets in public MCP configs in 2026, 2,117 of them still valid.
If a customer's security review is what's blocking the deal
Buy a named vendor — Lakera Guard/Check Point, Prisma AIRS or Cisco AI Defense — and map your controls to the OWASP Top 10 for Agentic Applications (Dec 9, 2025). The document is what enterprise reviewers are grading against; the filter is what they'll ask to see evidence for.

This analysis is drawn from the Guardrails element dossier — the ranked top 5, the comparison matrix, and the complete field of 30 more tools live there, with every source. Data: gd.json (CC BY 4.0).

Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.

Build your stack in 5 questions →