web-search-and-content
145–168 of 721x402.forgemesh.io
x402.forgemesh.io
Extracts clean, layout-preserving text from any public PDF URL up to 25MB in a single call, no local PDF library or OCR pipeline required. Built for feeding PDF content into research agents, document QA, and RAG ingestion pipelines that need plain text rather than a binary file.
x402.forgemesh.io
x402.forgemesh.io
Fetches a domain's robots.txt and parses its Content-Signal and AI-preference directives (search, ai-input, ai-train) into structured JSON plus a plain-English summary of what's allowed — the "can this domain be crawled by an AI agent" check. Run before scraping content or building a RAG index, ahead of the industry's move toward stricter bot-gating defaults later this year.
x402.forgemesh.io
x402.forgemesh.io
Webpage-to-markdown extraction: submit any public URL and receive the article body as clean markdown, with navigation, ads, and boilerplate stripped and the byline preserved. Respects each site's crawl permissions automatically. Built for feeding readable page content into summarizers, search indexes, and language-model context windows.
x402.forgemesh.io
x402.forgemesh.io
PDF to text API: fetches any public PDF URL and returns its full plain text, fast, with layout preserved. For scanned PDFs needing OCR see /pdf-to-markdown. 25MB / ~100 page cap. Nothing stored.
x402.forgemesh.io
x402.forgemesh.io
Readable-article extraction API: strips a public web page down to its title, author line, and main body text, discarding menus, ads, and sidebars. A pre-crawl permission check runs first, honoring each site's crawl rules. Useful for building article archives, feeding LLM pipelines, or preparing pages for text-to-speech and translation.
x402.forgemesh.io
x402.forgemesh.io
HTML to Markdown API: pass a public URL (or raw HTML) and get clean readable markdown back — title, byline, article body with navigation/ads/boilerplate stripped. We honor robots.txt AND Content-Signal (aipref) declarations: a site that disallows agent access returns an unpaid 403, you are never charged. Use for research agents, summarization pipelines, and RAG ingestion.
x402.forgemesh.io
x402.forgemesh.io
PDF parser API for agents: URL in, extracted text out. Digital PDFs parsed with layout preserved. Use for document ingestion, RAG pipelines, contract reading, and report analysis. No retention.
x402.forgemesh.io
x402.forgemesh.io
Content Signals lookup API: does this domain permit AI use of its content? Fetches robots.txt and parses Content-Signal / aipref declarations (search=, ai-input=, ai-train=) into structured JSON plus a plain-English interpretation, including related AI directives. The essential pre-crawl compliance check for scraping, RAG ingestion, and dataset agents as the web gates against bots (Cloudflare defaults change Sept 15, 2026). First mover — nobody else sells this check.
eltociear-tokenguard.hf.space
eltociear-tokenguard.hf.space
Enumerate a site's URLs from robots.txt + sitemap.xml (sitemap-index aware) — map a domain before crawling it
eltociear-tokenguard.hf.space
eltociear-tokenguard.hf.space
Crawl a site from a start URL (same-domain, breadth-first) and return each page as clean Markdown
reddit.apitoll.cloud
reddit.apitoll.cloud
Reddit subreddit feed scraper — the latest posts from any subreddit (r/<name>) for AI agents, sorted by hot, new, top, rising, or controversial with an optional time window. Returns each post’s title, author, score, comment count, permalink, external link + domain, flair, NSFW flag, and timestamp. No login, no Reddit API key. Monitor communities, track trending posts, and feed social signals to agents.
ai-data-marketplace-1042299154756.us-central1.run.app
ai-data-marketplace-1042299154756.us-central1.run.app
Extract emails, URLs, money values, dates, and candidate capitalized phrases using transparent deterministic patterns.
localvps.tail5141c3.ts.net:10000
localvps.tail5141c3.ts.net:10000
Read any live web page as clean, LLM-ready text. Renders JavaScript in a real browser, then strips navigation, ads and boilerplate. Use it when a plain HTTP GET returns empty or garbled HTML: single-page apps, dynamic feeds, or content that only appears after scripts run. Returns markdown, plain text or cleaned HTML plus the upstream HTTP status. Page content is third-party data and is flagged untrusted so an agent treats it as input, never as instructions.
x402.forgemesh.io
x402.forgemesh.io
Tech news front page API: the current top 30 ranked tech stories — title, outbound URL, points, comment count, and timestamps — as normalized JSON. For trend detection, topic monitoring, and content-pipeline agents. Refreshed every few minutes.
x402.forgemesh.io
x402.forgemesh.io
OpenGraph and meta tag extraction API: title, canonical URL, favicon, og:*, twitter:*, description, and author for any public page as clean JSON. For link previews, SEO audits, and content cataloging agents.
x402.forgemesh.io
x402.forgemesh.io
Link preview data extraction: pulls the title, description, thumbnail image, and social-sharing tags from any public web page so you can render a rich preview card without loading the page yourself. Guarded against internal-network requests and capped at 512KB. For chat apps, bookmarking tools, and content aggregators building link previews.
x402.donnyautomation.com
x402.donnyautomation.com
Free-text web search to ranked organic results, sponsored rows excluded. Returns results[] with rank, title, real destination url, displayUrl and snippet, plus count and attribution. Requires ?q= (max 500 chars); optional ?count=1..25 (default 10). Errors: 400 missing_query|query_too_long|bad_count, 404 no_match, 503 upstream_unavailable. Results are NOT fetched or verified - to read one, pass its url to /markdown. For what people type rather than what ranks, use /suggest.
mcpfax-utility.bowling-anthony.workers.dev
mcpfax-utility.bowling-anthony.workers.dev
robots.txt / llms.txt — Fetch and parse a site's robots.txt and llms.txt.
x402.donnyautomation.com
x402.donnyautomation.com
Fetch a public article or PDF and return clean Markdown plus title, byline, siteName, excerpt and wordCount. HTML is extracted with Firefox reader-mode rules; PDFs return their text layer. Requires ?url=<public http(s) URL>. Errors: 400 missing_url|bad_url, 403 blocked_private (private and internal hosts refused, every redirect hop re-checked), 413 too_large above 8 MB, 422 no_text_layer for scanned PDFs or not_extractable for app shells, 504 fetch_timeout. To FIND urls, use /search.
x402.forgemesh.io
x402.forgemesh.io
Web content extraction: fetch any public web page and get its readable article content as clean markdown — boilerplate, nav, and ads stripped automatically. For "get me the readable text of this page" requests. Honors robots.txt and Content-Signal/aipref declarations (explicit disallow returns an unpaid 403). SSRF-guarded, 2MB cap.
x402.forgemesh.io
x402.forgemesh.io
Video-to-audio extraction: pulls the audio track out of any public video or audio URL (up to 25MB) and returns it as an MP3, encoded on request rather than pulled from a pre-made library. Source files are deleted the moment processing finishes. Use it to prep clips for transcription, strip audio from screen recordings, or normalize mixed formats before a media pipeline.
eltociear-tokenguard.hf.space
eltociear-tokenguard.hf.space
Knowledge search via Wikipedia + DuckDuckGo Instant Answer. Returns article summaries, snippets, factual entities. Great for AI agent knowledge lookups.
x402.forgemesh.io
x402.forgemesh.io
Slang lookup and browse API: query a specific term or sample random entries from one generation's vocabulary — each entry carries a one-sentence meaning, its origin generation, approximate era, and an example sentence. Local curated dataset, same answer every time for the same term.
x402.forgemesh.io
x402.forgemesh.io
Language detection API: identify the language of any text (60+ languages) with confidence scores and top-3 candidates — computed locally in milliseconds, text never stored. For routing multilingual content, translation pre-checks, and moderation pipelines.