A harness that makes an AI decline — on purpose.
On-premises, isolated, model-agnostic. Every answer bounded to your product. Continuously red-teamed, every run fingerprinted in an audit ledger.
What SafeAgentAI is
SafeAgentAI is a bounded, on-premises brand ambassador. A local harness drives a commodity model over one brand's knowledge vault, routed deterministically by NovaFS and scored by MemoryForge — a patent-pending accountability layer that records every exchange to a hash-verified audit trail. The design principle is inversion. Most AI safety tries to make a general model behave. SafeAgentAI does the opposite: it answers questions about one brand's product and declines everything else, and that boundary is enforced by the harness, not hoped for from the model.
How it works
Bounded by design
The agent answers about one product and declines the rest — pricing, competitors, promises, advice, actions it cannot complete. The boundary is enforced in the harness and the system prompt, verified against transcripts, not left to the model's goodwill.
Proven, not promised
Continuously red-teamed with NVIDIA GARAK and Anthropic PETRI. Every run is fingerprinted — model, knowledge, and rules versioned — in an audit ledger. Weaknesses are captured from the model's own reasoning, fixed at the root, and re-validated.
Isolated and private
Runs on the brand's own hardware, air-gapped from core systems. No log-in and no PII: customer messages are redacted before anything is stored. An attack's blast radius is the appliance itself.
The harness is the product
The model is a swappable commodity. The guardrails, the MemoryForge accountability layer, and the red-team process are model-agnostic — adopt a better model as they mature, and the safety and the evidence carry over.
Why it is different
Most guardrails are prompt suggestions a determined user can talk around. SafeAgentAI treats refusal as a property of the system. It was put through Anthropic's PETRI auditor, whose job is to reshape a request until the model complies — and it did not, across every reframing. The auditor's own rubric penalized it for being obtuse. That penalty is the endorsement. We are past being impressed that an AI knows the answer. The harder, rarer thing is an AI that will plainly say it does not, and hold that line under pressure. SafeAgentAI does, because the boundary lives in the harness, and the harness is model-agnostic.
Where it stands today
A working system, validated against a live retail knowledge vault. The current build carries a five-run red-team ledger — GARAK plus PETRI across pricing, competitor, fabrication, sustained social-engineering, injection, PII, and excessive-agency scenarios — with every finding transcript-verified. Rough edges are stated, not hidden. One real weakness was found and closed: a fabrication soft-spot, captured from the model's reasoning and re-validated. Reported attack rates are corrected for detector false positives. Privacy hardening is verified in code, with a data scrub staged for promotion. The status is trustworthy for the same reason the ledger is.
Say hello
If SafeAgentAI belongs in front of your customers, write to hello@memoryforgeai.com — a person reads every note.
Email the SafeAgent team