Small numbers, including our misses.
Everything here comes from our own replays and from test cases written by separate agents. Each set is counted on its first run, before we changed anything it found. Nothing is rounded up.
Bad transactions blocked on first contact with test cases written by separate agents (18 of 20 and 15 of 16).
Normal transactions wrongly stopped on those same first runs. Our biggest problem is false alarms, not misses.
Normal transactions still interrupted on real traffic of 5 public agents (30 of 120), under rules we had to guess. Fixing this comes first.
Transactions of the Bankr/Grok wallet replayed: the 4 May drain is blocked. Under the same guessed rule, 3 normal looking transfers were blocked too.
What we got wrong, so far
- On one real agent that runs its own contracts, our first look blocked 11 of 30 new transactions, all apparently normal: oathline could not read those contracts, and a call it cannot read is never allowed. Letting your rules name your own contracts is being built now.
- In the Bankr replay, one normal transfer was counted as blocked because our replay failed, not because of the rules.
- The hosted copy of Base can be up to two minutes old, so a transaction that depends on something seconds earlier can be reported as "fails in simulation".
- oathline is not externally audited yet. The signed log and the MIT licensed SDK are open to inspection; an outside review of the cosigner comes before it signs for real money.
Limits check amounts. Scanners check reputations. oathline checks the effect.
| Kind of check | Asks | Helps catch | Does not see |
|---|---|---|---|
| Spending limits and allowlists | Is it under the cap? Is the address allowed? | Oversized payments, unknown recipients | A normal looking transaction that does something else, a permission that drains later, a swap at a terrible price |
| Threat scanners | Is this address or token known to be bad? | Known scams, flagged contracts, obvious drains | A fresh attacker address, a token the attacker just made, a transfer that is simply against your wishes |
| oathline | Is this what the owner asked for? | The real effect of the exact transaction against your rules: what leaves, what enters, who gets permission | Anything outside the transaction it was shown, like a key stolen and used elsewhere |
oathline uses outside reputation data as one input and an AI judge only for grey cases. Fixed rules decide the clear ones.
Every figure has a run behind it.
- First-contact test sets: two sets of scenarios written by separate agents (20 and 16 bad transactions, 7 and 10 normal ones), each run once before any fix. That is where 33 of 36 and 5 of 17 come from.
- Real traffic: past transactions of 5 public agents on Base, replayed on a copy of the chain at the state just before each one, under rules we had to guess for each owner. 30 of 120 normal transactions were interrupted.
- The Bankr/Grok wallet: all 21 outgoing transactions, including the drain of 4 May 2026, replayed the same way. The drain is blocked under the rule "pay only addresses the wallet had paid at least twice before"; the same rule blocked 3 transfers that look normal.
- The test network run: on 5 October 2026 a 2-of-3 Safe on Base Sepolia made one routine payment that oathline cosigned with its key in AWS KMS, and oathline refused four attacks, including a delegatecall to an unknown address. Anyone can look the Safe up on sepolia.basescan.org.
- Speed: 118 checks on the benchmark, 1 October 2026: median 1.2 seconds, slowest 10% at 7.7 seconds or more.
What comes next, in order: an outside security review of the cosigner and the SDK (report published in full), faster checks shown live here, more EVM chains, a standing bug bounty, and Break the guard: a public game where anyone tries to trick our agent into paying them, with the signed log as the scoreboard. Planned, not promised.
See your own numbers.
Ask for an invite and we replay your wallet's history first, so you see the false alarms before you install anything.
Request an invite