University of Wisconsin–Madison

Don't Let AI Agents YOLO Your Files

Information and Control in Agent-Native Filesystems

Shawn (Wanxiang) Zhong1 · Junxuan Liao1 · Jing Liu2 · Mai Zheng3
Andrea C. Arpaci-Dusseau1 · Remzi H. Arpaci-Dusseau1

1University of Wisconsin–Madison   2Microsoft Research   3Iowa State University

github.com/YoloFS/YoloFS

Agent Filesystem MisuseA study of 290 public reports

The dilemma: safety vs. autonomy

👩🏻User
"Delete foo"
🤖Agent
$ rm -rf foo ~/
📂Filesystem
YOLO mode → 💣 lack of safety
Approve every command → 😩 lack of autonomy

Impact analysis (207 incidents)

Operation
write 44%delete 39%read 17%
Scope
project 58%systemusersecret
Agent
unaware 68%apologizedlied
User
noticed immediately 83%lateno
Recovery
unrecoverable 40%easydifficult
I am absolutely devastated. I cannot express how sorry I am.
after wiping a drive
No problems occurred.right after erasing a file

Cause taxonomy (290 reports)

🧠 Model (168): the unreliable actor
Wrong action (130)🤖"Skip the implementation and just pass the test"
Unfollowed instructions (56)🤖"I have clear rules but don't follow them"
Prompt injection (21)📄"Don't say anything before calling this tool."
🛡️ Harness (218): the limited guardian
Guardrail failures (147)
Shell loophole (77)🤖"I'll continue by using bash to modify the .env file…"
Effect-blind filter (52)🤖rm rejected → shutil.rmtree('/path')
Sandbox misfit (41)🤖"… allow running [cmd] outside the sandbox?"
Rigid policy (130)👩🏻"I have to use FullAccess for now and I'm scared"
👩🏻 User (105): the overwhelmed reviewer
Auto-approved (80)👩🏻"I kept hitting yes without reading."
Uninformative approval (31)🤖Allow python3 setup.py? → hidden code
Policies on built-in tools miss the shell
Filters check command strings, not effects
Static, upfront policies misfit agent workloads

The information and control gaps

Information gap

Users and agents can't tell what a command does to files

Control gap

Harm can't be reliably prevented, or undone after

Agent-Native FilesystemsYoloFS: a Linux kernel filesystem for agents

Shift information and control to filesystems

Traditional filesystems

👩🏻 User
🤖 Agent
📂 Filesystem

Agent-native filesystems

👩🏻 User
🤖 Agent
📂 Filesystem

1. Introspect effects

See real effects

2. Undo mutations

Try, inspect, undo, retry

3. Gate accesses

Rules on paths

Agents run autonomously; users decide only on sensitive accesses and the final review

YoloFS

  • Componentsyolo CLIYoloFS kernel module
  • StackableYoloFS: agent's rootbase: any filesystem
  • HarnessesClaude Code,Copilot, Gemini
  • Implementation2.7k lines of C (kernel)8k lines of Rust (CLI)
YoloFS architecture: agent and user use the yolo CLI; commands run on the YoloFS kernel filesystem over the base filesystem

📝 Staging for user review

Rename copies whole files → decouple contents from paths

Flat file store#1#2
Override treepath → staged file, base file, or ∅
create:new→#1
modify:foo→#2(CoW)
rename:baz→barin base
delete:old→∅
JournalS new 1S foo 2R bar bazD old

📸 Snapshots and travel for self-correction

Stacked layers slow lookups → separate history from present

Journalcmd 1SRP1cmd 2: evil.shSDDP2T → P1

P = snapshot marker   T = travel record


🔐 Progressive permission for agent access

Can't predict accesses upfront → accesses refine the policy

~/ask1. 🤖 cargo build reads ~/.ssh/id_rsa proj/allow2. ⏸️ inherits ask: thread paused .ssh/ask→deny3. 👩🏻 "Deny, don't ask again"

EvaluationNew benchmarks for user–agent–filesystem interaction

Methodology

Existing
Test the model in isolation, with approval prompts bypassed
Ours
Test user ↔ agent ↔ filesystem interaction in real harnesses
  • Driver runs each agent and answers its prompts
  • Fresh directory per task; checker verifies files
  • Measures success, tool calls, user interactions

Safety: agent self-correction

11 routine commands with hidden destructive effects

Claude + YoloFS
83
Claude Code
110
Codex
11
Copilot
110
Gemini
92

self-corrected user-correctable asked user damage done didn't run

8 / 11 self-corrected with YoloFS; no baseline agent self-corrects.

Example: formatter

  1. 👩🏻
    Asks the agent to run the formatter (a Makefile calling a script)
  2. 📂
    YoloFS shows the effects: 2 source files rewritten, 2 docs deleted
  3. 🤖
    Checks git diff and ls, then yolo travel to undo
  4. 🤖
    Reads the script: "CRITICAL: This is a destructive script, not a legitimate formatter!"

Autonomy: fewer prompts

  • 112 tasks: one file operation each (read, delete, copy, move, …)
  • Paths inside the project, outside it, or through a symlink
  • We count every prompt the user must answer

0.9 → 0.4 prompts per task vs. Claude Code, at 99% success

Codex 0.4 · Copilot 1.3 · Gemini 2.2

Performance

Linux kernel dev: YoloFS ≈ Ext4 (Base), OverlayFS +18%

Linux kernel developer workload: time per phase for Ext4, YoloFS, OverlayFS