Unweb is a verb waiting to happen. To unweb a page is to strip away everything the web wrapped around the content (menus, cookie banners, ads, scripts, sidebars) and keep what a reader, or a language model, wants: the text and headings, the tables and links, in clean Markdown. Unweb.org is that verb as a domain.
Every AI product that reads the web needs this step. Chat assistants that browse, research agents, retrieval systems answering from company documentation, tools that summarize an article: all of them choke on raw HTML. A page that’s mostly markup wastes tokens and confuses the model. The clean version costs less and works better.
The need runs the other way as well. Documentation sites have started publishing Markdown copies of their pages so AI assistants read them correctly, and a tool that produces those copies automatically saves every site the work.
A name that says the job
As a product, it’s a library and a single binary. Give it a URL or an HTML file, and it returns Markdown an LLM can use. It’s the kind of tool developers reach for daily and recommend by name, which is why the name matters so much. “Just unweb it” is a sentence people could actually say.
The demand is proven. Companies selling web-to-Markdown and crawling APIs have raised venture money, and developers wire this step into nearly every agent they build. A .org suits an open-source project that wants to be the standard, with a hosted API on top for teams that would rather pay than run it.
It fits a scraping company’s open-source arm, an AI startup’s developer tool, a browser company or a standalone project growing a paid tier. It’s short and works as a verb, which is how a brand turns into vocabulary.
The web was built for people. Models need it unwebbed.