Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori
To learn what a US school district is buying, you file a Freedom of Information Act request. Someone scans the email you sent, puts the scan on Google Drive, and attaches the relevant PDFs. Dhruv Batra's question is whether anyone seriously expects that office to publish an MCP server. He grants the popular claim that agents will drive most of the action on the web, then rejects its usual next step, that the web will meet them with APIs. The head of the distribution might. The long tail, some 200 million active sites where infrastructure changes over decades, will not.Reading the HTML instead does not save you, because much of what you see was never written down anywhere. A basketball score is missing from the page that first loads and arrives later as JSON. A product page contains no text reading sold out, only a quantity of zero and a script that grays the option out. State is calculated and rendered rather than stored, which makes the browser closer to a game engine than a document and makes pixels the source of truth. He calls this the bitter lesson for web agents: scaffolding built per site does not generalize, and the general solution is the one the web was actually built for. Their Navigator model runs screenshot in and clicks out, now writes JavaScript when that is quicker, and checks the result on screen. It misses 8 of 300 trajectories on a benchmark he thinks should be retired.
Speaker info:
- https://x.com/DhruvBatra_
- https://www.linkedin.com/in/dhruv-batra-dbatra/
- https://dhruvbatra.com
Timestamps:
0:00 - The argument, and the part of it that is wrong
3:23 - Restaurant menus on easy, medium, and hard mode
5:31 - School district procurement, up to a FOIA request
7:38 - 200 million active sites that change slowly
8:42 - Why reading the HTML does not rescue it
10:14 - Sold out is not text, it is a rendered zero
11:33 - The browser is a rendering engine, pixels are the truth
12:53 - Navigator, and writing JavaScript when that is faster
15:51 - Are computer use models actually good enough
17:20 - Latency and cost per task
18:28 - Another layer of mess that we will call an API Receive SMS online on sms24.me
TubeReader video aggregator is a website that collects and organizes online videos from the YouTube source. Video aggregation is done for different purposes, and TubeReader take different approaches to achieve their purpose.
Our try to collect videos of high quality or interest for visitors to view; the collection may be made by editors or may be based on community votes.
Another method is to base the collection on those videos most viewed, either at the aggregator site or at various popular video hosting sites.
TubeReader site exists to allow users to collect their own sets of videos, for personal use as well as for browsing and viewing by others; TubeReader can develop online communities around video sharing.
Our site allow users to create a personalized video playlist, for personal use as well as for browsing and viewing by others.
@YouTubeReaderBot allows you to subscribe to Youtube channels.
By using @YouTubeReaderBot Bot you agree with YouTube Terms of Service.
Use the @YouTubeReaderBot telegram bot to be the first to be notified when new videos are released on your favorite channels.
Look for new videos or channels and share them with your friends.
You can start using our bot from this video, subscribe now to Computer-use models will agentify the web, not APIs — Dhruv Batra, Yutori
What is YouTube?
YouTube is a free video sharing website that makes it easy to watch online videos. You can even create and upload your own videos to share with others. Originally created in 2005, YouTube is now one of the most popular sites on the Web, with visitors watching around 6 billion hours of video every month.