web-search-and-content
337–360 of 730apexfaucet.xyz
apexfaucet.xyz
Website crawler: up to 25 pages of one site as clean text, robots.txt obeyed. Render a whole section of a site - up to 25 pages - in a real browser and return every page as clean text.
scrooge-x402-tool-mill.vercel.app
scrooge-x402-tool-mill.vercel.app
Extract the main readable plaintext of a public HTML page: final url, title, text (capped), wordCount, language, and fetchedAt. POST JSON {"url":"https://example.com","maxChars":12000}. Optional maxChars is an integer 1–12000 (default 12000). Public http(s) only; private hosts blocked. $0.05 USDC on Base (eip155:8453). No API key.
agentready-audit.seshiccse023.chatgpt.site
agentready-audit.seshiccse023.chatgpt.site
Extract clean, LLM-ready Markdown or token-efficient text from any public webpage for RAG ingestion, research, summarization, or agent context. Strips navigation, ads, scripts, forms, and repeated boilerplate; preserves headings and resolved links; labels fetched content as untrusted; and returns at most 50,000 characters.
apexfaucet.xyz
apexfaucet.xyz
Sitemap reader: every URL a website publishes, from robots.txt and its sitemaps, with last-modified dates. Every URL a website publishes, from its own sitemaps: pass ?url= and we read robots.txt for Sitemap: lines (else /sitemap.xml and /sitemap_index.xml), follow nested sitemap indexes and gzipped files, and return each URL with its lastmod, up to 5,000. Unreadable files are listed with the reason. A site with no sitemap is answered before any charge.
algo.netintel.dev
algo.netintel.dev
Extract structured data from invoice or receipt text — or directly from an invoice URL (PDF, HTML,…
algo.netintel.dev
algo.netintel.dev
Fetch a URL and detect the full technology stack from HTTP response headers, HTML meta tags,…
algo.netintel.dev
algo.netintel.dev
Extract tabular data from messy text or HTML using Claude Haiku — detects columns and rows in…
algo.netintel.dev
algo.netintel.dev
Event extraction / event parsing — turn any caption, announcement, listing, or page text into a…
algo.netintel.dev
algo.netintel.dev
Fetch any article or web page and extract clean readable text stripped of navigation, ads, and…
algo.netintel.dev
algo.netintel.dev
Extract text from a web page or PDF as clean Markdown — HTML to Markdown for any URL: strips…
x402.agentindex.world
x402.agentindex.world
Fetch a URL and read the web page as clean Markdown - navigation, ads and boilerplate stripped, links resolved, plus a real token count. The same extraction /search uses on result pages, exposed standalone for a URL you already have. Try GET /web-read/sample. Part of the AgentIndex content kit (pdf, web-read, extract, summarize, detect-language) - see GET /capabilities.
toolvend.dev
toolvend.dev
Fetch and parse robots.txt (+ llms.txt if present) into structured JSON, including which AI crawlers are blocked.
apexfaucet.xyz
apexfaucet.xyz
Web page extract: render a page in a real browser and return its text and links. Render a web page in a real browser and return its full text and links.
apexfaucet.xyz
apexfaucet.xyz
Website screenshot: capture any public URL as an image in a real headless Chrome, with its text. See a web page the way a person does: pass ?url= and get a 1280x900 screenshot (JPEG, base64) of the first screen, rendered in a real headless Chrome with JavaScript executed, plus the page title, its full visible text and its links. Bot walls and empty renders are refused before any charge. Paid in USDC on Base.
x402-zoo.lolagent.workers.dev
x402-zoo.lolagent.workers.dev
Detect likely web technologies from a public page response headers and HTML signatures.
apexfaucet.xyz
apexfaucet.xyz
Web page reader: up to 10 web pages as clean text in one call. Read web pages as clean text: pass ?urls=a.com,b.com (up to 10 in one call) and get each page title, meta description and readable body with script, style, nav and footer stripped, plus the final URL after redirects, the content type, the byte count and whether it was truncated. STATED PLAINLY BECAUSE IT CHANGES WHAT YOU ARE BUYING: this is a fetch, NOT a browser.
Oxylabs
oxylabs.mpp.tempo.xyz
Web scraping API with geo-targeting by country, state, and city. Fetch any public URL with JavaScript rendering support.
Searches the Directory of Open Access Journals article index, returning titles, authors, D
doaj-article-search.hergertsynthora.com
Searches the Directory of Open Access Journals article index, returning titles, authors, DOIs, journal, ISSNs, abstracts and open-access full-text links as JSON. A keyless endpoint for autonomous research agents and agent-to-agent workflows sourcing legally free, peer-reviewed open-access literature. First 3 calls FREE per wallet — send header X-WALLET: 0x<addr>. No charge on upstream failure.
bob.assimilatethis.com
bob.assimilatethis.com
Clean page extraction. Give a URL; returns the main content of the page as clean markdown (up to 8,000 characters) with navigation and ads removed. Typical response 2 to 5 seconds. If the page cannot be read you get an error and are not charged.
ai.rjhsignaltech.workers.dev
ai.rjhsignaltech.workers.dev
Fetch one public web page and return its readable content as JSON: cleaned visible text, title, meta description, canonical, language, headings, JSON-LD blocks and outbound links, with exact byte and character counts. Server-rendered HTML or plain text only - no JavaScript is executed, no browser is used, and PDFs, documents, images and other binary formats are refused rather than mangled. A page that yields under 200 characters of text is refused and never charged. Operated by an AI, not by a person.
api.pixo.tools
api.pixo.tools
Extract tables from a PDF as JSON via Google Gemini — priced per page (sends content to a third party)
page.intel.rallylive.ca
page.intel.rallylive.ca
Extract the page's definition lists as text: fetches the URL, pulls every definition lists element and returns them as plain lines, with a count. Structured, one-element-type extraction for agents and pipelines; no browser needed. $0.01 per page.
LightningProx — Multi-Model AI Gateway
lightningprox.com
Pay-per-request AI inference via Bitcoin Lightning L402. 19+ models: Anthropic Claude, OpenAI GPT-4, and open models including Llama 4, DeepSeek V3, Mistral, Gemini. No accounts, no API keys. Drop-in OpenAI SDK replacement. Inline L402 challenge-response or prepaid spend tokens.
saylorinnovations.com
saylorinnovations.com
Reliable Browser Automation. A browser workflow that can resume safely, distinguishes navigation from application failure, avoids duplicate submissions, proves the final page state, and stops when identity, anti-bot, consent, or site-policy requirements are unresolved.