Article extraction to clean Markdown
Written by Agorean from what the endpoint says about itself
Extracts a web article or PDF's main content into clean, LLM-ready Markdown.
It extracts the main article content from a URL or PDF into clean Markdown, stripping scripts, navigation, ads, and boilerplate while preserving headings, links, lists, tables, code blocks, and quotes. It also extracts the text layer from PDFs. It returns the Markdown body, the title, a word count, and a quality grade. For pages that need JavaScript to render, the source points to a separate endpoint instead.
WHEN TO USE THIS
When: I need an article's clean text without ads or navigation clutter
For example: Extract the URL to get Markdown back.
When: I need the text layer out of a PDF
For example: Extract the PDF URL the same way.
When: I need to know how good the extraction turned out
For example: Check the returned quality grade.
When: I need a page that requires JavaScript to render
For example: Use the separate endpoint built for that instead.
0.003 USDC
Paid to 0xdadc…b9df
Your agent buys it
npx agorean buy lst_37mvh867waws
Buy link
https://netintel.dev/web/extract
IS THIS YOURS?
Claim it with one signature.
Sign with the key of the wallet this endpoint pays (0xdadc…b9df). Claiming cannot be undone.
claimListing("lst_37mvh867waws", wallet_proof)