Scraper policy.

Last updated 18 August 2026

This is the page our crawlers point to when they visit your servers. It says who we are, what we fetch, how our crawlers behave, and how to reach us. What happens to the data after we fetch it is covered by our data use page.

Who we are

Largesso is an art-discovery platform. We index public collections from museums, galleries, and art platforms so that people can find works and follow them back to the institutions that hold them. Every work we display carries a credit line and a link to the source institution's own page for it.

How our crawlers identify themselves

Traffic from us arrives under one of these User-Agent strings:

largesso-pipeline/1.0 for catalog metadata, and Mozilla/5.0 (compatible; LargessoBot/1.0; +https://largesso.com) for image fetching. Both carry a contact address. We do not disguise our crawlers as browsers, and traffic from us never arrives under a spoofed identity.

How they behave

Where an institution publishes an API, bulk export, or data dump, we use it instead of crawling pages. Where an institution documents a rate limit, that limit is written into our fetcher as a hard per-host floor it cannot cross. Hosts without a documented limit get a conservative default rate and a small connection cap.

If a server blocks us, we stop and ask. We do not rotate proxies or otherwise disguise traffic to get around a block, and a source that has asked us to stop stays off the platform until it gives explicit permission.

What we fetch

Catalog metadata (titles, artists, dates, credit lines) and preview images at modest resolution. How that material is hosted, embedded, attributed, and taken down is described on the data use page, and the per-source state is public at largesso.com/sources.

Contact and opt-out

Questions about our crawling, a rate you would prefer we use, or a request to stop entirely: email [email protected] or reach us through support. One email from the institution removes the source the same day; no legal letter required.