Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

A script under one megabyte that never looks at the screen matches or beats the frontier model it was copied from. Pierluca D'Oro builds it by recording one successful trajectory per task and then replaying those actions blindly, and on deterministic benchmarks like OSWorld that counts as a valid agent and it scores at the top. The paper goes further and proves that pass@k on a deterministic environment is exactly the success rate of that replay script, so a metric the field leans on turns out to be a formal measure of the exploit.

The fix has two halves. Environments get the PRISM principles: privileged verification, realism, integrity checked configurations, sandboxed execution, and multifactorial variation across data, theme, and starting screen. DIGIWORLD instantiates them in 15 sandboxed mobile apps and 3.2 million verified configurations, generated by a compiler that produces every combination and rejects the broken ones, because a coding agent emitting a lot of software is not the same thing as a good environment. Metrics get honest uncertainty. Naive rollouts on a single base case yield confidence intervals that actually contain the true performance around 20% of the time rather than 95%, and he prices the consequence: a 4% gap between two models, hidden under intervals that look tight, costs hundreds of thousands of dollars a month across a million tasks.

Speaker info:
- https://x.com/proceduralia
- https://www.linkedin.com/in/pierluca-doro/
- https://www.proceduralia.com
- https://arxiv.org/abs/2605.08261

Timestamps:
0:00 - The replay agent, a script that never sees the screen
1:41 - It matches the model it was copied from
2:22 - Why pass@k measures exactly that exploit
3:50 - The PRISM principles for environment design
5:31 - DIGIWORLD, 15 apps and 3.2 million verified configs
7:25 - The compiler that rejects invalid combinations
9:21 - Replay stops working, and frontier models look fragile
11:05 - Two sources of variance, actions and environment
12:31 - Intervals that cover 20% of the time, not 95%
14:56 - A benchmark without rigor is a misleading one Receive SMS online on sms24.me

TubeReader video aggregator is a website that collects and organizes online videos from the YouTube source. Video aggregation is done for different purposes, and TubeReader take different approaches to achieve their purpose.

Our try to collect videos of high quality or interest for visitors to view; the collection may be made by editors or may be based on community votes.

Another method is to base the collection on those videos most viewed, either at the aggregator site or at various popular video hosting sites.

TubeReader site exists to allow users to collect their own sets of videos, for personal use as well as for browsing and viewing by others; TubeReader can develop online communities around video sharing.

Our site allow users to create a personalized video playlist, for personal use as well as for browsing and viewing by others.

@YouTubeReaderBot allows you to subscribe to Youtube channels.

By using @YouTubeReaderBot Bot you agree with YouTube Terms of Service.

Use the @YouTubeReaderBot telegram bot to be the first to be notified when new videos are released on your favorite channels.

Look for new videos or channels and share them with your friends.

You can start using our bot from this video, subscribe now to Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

What is YouTube?

YouTube is a free video sharing website that makes it easy to watch online videos. You can even create and upload your own videos to share with others. Originally created in 2005, YouTube is now one of the most popular sites on the Web, with visitors watching around 6 billion hours of video every month.