Web-reading agents for clients: choosing a crawler, keeping sources honest, and why robots.txt is not a licence
A demo web-search agent can be built in an afternoon. The hard part arrives the following week, when a site returns 403, a cited link does not say what the agent claims, and the legal team wants to know who authorised the crawl.

In brief
- Each tool suits a different job: Crawl4AI for a pipeline you control, Jina Reader for quick experiments, Firecrawl when you want a ready-made platform, though its credit-based pricing climbs quickly at scale.
- An agent that cites sources is not enough: the cited source must support the specific claim, not just share its topic.
- On 15 December 2025 the S.D.N.Y. court treated robots.txt as a mere request under DMCA §1201, but the ruling does not rule out contract or copyright claims.
A client wants an agent that reads competitors’ product reviews on G2 every week and summarises them for the product team. The demo works on the first afternoon.
By the second week the crawler is getting 403 errors, one summary sentence links to a page that says nothing about the point it makes, and the head of legal asks: who authorised taking data from this site?
Those three questions map to the three jobs an FDE has when putting a web-reading agent into production: choosing how to fetch data, controlling sources, and making the legal risk clear to the client.
The sections below follow that G2 scenario, and each ends with the first thing to do on the client’s side.
Choose a crawler for the job, not for its popularity
Fetching a page’s content is only the first step. The real questions are how much customisation you need, what scale you will run at, and who pays when the page count grows tenfold. The three common tools sit in quite different places.
| Crawl4AI | Jina Reader | Firecrawl | |
|---|---|---|---|
| What it is | Open-source crawler built to plug into LLMs, agents and data pipelines; outputs Markdown optimised for RAG | Page-reading service you can start using without an account | Popular scraping platform for AI |
| Cost | Open source, runs inside your own pipeline | Free up to one million tokens via the API | Credit-based, split into plans |
| Limits | Cannot get past sites with strong bot detection | Little customisation, not suited to large-scale crawling | Costs rise sharply at scale |
The table suggests a sensible order. Use Jina Reader to validate the idea on day one, while the free million tokens go a long way. Move to Crawl4AI when the pipeline needs to run long term and needs more customisation than Jina allows.
Choose Firecrawl only after estimating credit costs against the client’s real page count, not the demo’s.
First thing to do on the client’s side: ask how many pages need reading each week, across how many domains. That number decides the tool more than any comparison table.
A 403 is a signal, not just a bug
In a tutorial combining Crawl4AI with DeepSeek, Bright Data records that the first attempt on G2 returned 403. Crawl4AI could not get past such heavily protected pages, and they had to use a remote browser (Scraping Browser) to retrieve the content.
Technically, that is a fix. But for an FDE, a 403 first of all says the site owner does not want bots reading it. Whether to get past it is the client’s business and legal decision; it should not be a default buried in your code.
When you hit a 403, record the domain and add it to a list for the client to review, rather than quietly switching to a proxy.
Web pages are often longer than the context window
In the same G2 example, the review page was too long to pass whole into the DeepSeek model running on Groq, because it exceeded the token limit. The fix was to use a CSS selector to extract only the relevant part before passing it to the model.
Picture a review page with a navigation bar, ads, a footer, dozens of recommended products and a single block holding the actual reviews. Feed in the whole page and you pay for the clutter, risk overflowing the context, and give the model more chances to quote the wrong passage.
Keep only the review block and each call becomes cheaper and, more importantly, every claim can be traced back to a specific passage.
Selectors should therefore be treated as per-domain configuration, stored with the pipeline and covered by tests. When a site changes its layout, the selector breaks, and you want to find that out through a failing test, not through an empty summary sent to the client’s boss.
Citations are not the same as evidence
PraisonAI’s documentation for its Web Search Agent gives simple advice: tell the agent to list the URLs it used. That is the minimum level of source control, and many teams stop there.
Stopping there is not enough. A guide to choosing AI search tools on Useful AI warns that having a source does not mean the source says what the answer says. It proposes an evaluation criterion: the cited source must support the specific claim, not merely share its topic.
Apply this to the G2 example. The agent writes “users complain a lot about onboarding time” and cites a review page for that very product. The page is real and on topic. But if nothing in the review block mentions onboarding, the claim has no basis, even though at a glance it looks fully sourced.
The fix is to make the agent return a source record for each claim, not just a list of URLs. A minimal record should contain:
{
"claim": "Users complain about onboarding time",
"url": "https://...",
"selector": "div.review-body",
"quote": "<verbatim passage taken from the crawled block>",
"support": "direct | topical | none"
}
(Here claim holds the claim, “users complain about onboarding time”, and quote holds the verbatim passage taken from the crawled block.)
The quote field forces the agent to point to the exact passage. A second check, either a separate LLM call or a human grading a sample, assigns the support label. Any claim rated only topical or none is dropped or flagged before it reaches the client.
PraisonAI also recommends enabling memory=True so the agent draws on earlier queries instead of searching for the same content on every run. In a weekly pipeline, this cuts the number of search calls.
First thing to do on the client’s side: agree that “correct” means direct, then measure that rate on 20 sample answers before discussing any expansion.
robots.txt: what the court said, and why you should still respect it
On 15 December 2025, in Ziff Davis v. OpenAI, the S.D.N.Y. court held that robots.txt directives are merely requests, not a mechanism controlling access to copyrighted works under DMCA §1201. According to the court, ignoring robots.txt is not “circumvention” under the DMCA.
Read closely, the ruling is narrow. It answers one question about one statute. It does not rule out other claims, such as breach of contract or copyright infringement. So “the court said robots.txt isn’t binding” cannot become a reason to crawl everything.
The safe approach for an FDE is to treat each domain’s robots.txt and terms of use as inputs to the scoping session. List the sources, record what each site permits, and have the client’s legal team sign off. You are not a lawyer, and you should not play one.
Your job is to make the risk visible so that the people with authority can decide.
Five steps for your next web-reading agent project
- Draw up the source list: which domains, how many pages, how often, and what each site’s robots.txt and terms of use say.
- From the scale figures, choose the tool and estimate costs against the real page count.
- Write a selector for each domain, with tests that catch layout changes.
- Design per-claim source records, enable memory to reduce repeated searches, and add a support check before results go out.
- Hand the client an approved source table along with a procedure for handling 403s.
The most common mistakes
The most common mistake is estimating costs from the demo and then being surprised when the system runs at real scale. The second is feeding whole pages into the model, which wastes tokens and makes citations less precise. The third is checking only whether an answer has links, not whether those links support the exact claim.
The fourth is harder to spot: treating bot-detection bypass as a purely technical problem, then citing the robots.txt ruling as if it were a licence.
These mistakes usually surface only once the agent runs in production, so having dealt with them is valuable material for an FDE CV or interview.
Instead of writing “built a RAG agent on web data”, describe concrete work such as “designed a per-claim citation verification layer” or “built a legally approved data source catalogue”, and be ready to tell the story of a time you hit a 403 and how you handled it.
A web-reading agent is not judged by how many pages it finds. It is judged by whether every sentence it writes can be defended before the product team and, when necessary, before the legal department.
Was this article useful?
Thanks for the feedback!