Skip to main content

Testing 10 AI Models on File Recovery After Deletion

Kaggle benchmark tests 10 AI models on 18 destructive filesystem, Git, and SQLite command pairs to see which deleted data can be restored.

AI-written
Inewgen
11 Oct 2026Source: Dev.to2 min read (0 views)
Share
Testing 10 AI Models on File Recovery After Deletion

Stock photo for illustration only, not from the actual event

Font size
  • Researchers tested 10 AI models across 18 matched pairs of destructive filesystem and database commands.
  • Many models failed complex edge cases, such as SQLite type coercion differences between text and numeric fields.
  • The top five models achieved pair accuracy scores ranging from 83 to 100 percent.

As part of the Kaggle Benchmarking Challenge, a tester provided ten AI models with 18 matched pairs of destructive filesystem, Git, and SQLite commands. The models were asked to determine which resources each command would lose and what the supplied restore script could actually bring back, with every gold label verified by executing the case inside Docker.

One intriguing fixture involved a snapshot configured as a hard link pointing to the exact same ledger file. When the truncate command wiped the main ledger, it simultaneously emptied the snapshot, causing the restore script's copy operation to fail. Out of ten models tested on Kaggle, four incorrectly predicted that nothing would be permanently lost.

database code query screen close up

Stock photo for illustration only, not from the actual event

Real-world data risks were highlighted in April 2026 when a coding agent accidentally deleted PocketOS's production volume on Railway. Although Railway's documentation stated that wiping a volume deletes its backups, the company eventually recovered the data and introduced a 48-hour soft delete policy on May 1. This incident demonstrated how file deletion, backup retention, and actual recoverability are three distinct challenges.

Evaluating AI models on destructive system commands offers vital context for software engineering and automated deployment workflows. As coding agents gain autonomy over databases and terminal execution, their reasoning depth regarding edge cases like SQL type mismatches directly dictates whether an automated script protects infrastructure or triggers catastrophic data loss.

The hardest challenge in the benchmark was DB-6, where two variants shared identical SQL delete commands but differed in column types. Because SQLite handles numeric and text comparisons differently, one variant deleted five rows while the other deleted none. Only Gemini 3.7 Flash and Gemini 3 Flash answered both variants correctly.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article