Scraper policy.
Last updated 18 August 2026
This is the page our crawlers point to when they visit your servers. It says who we are, what we fetch, how our crawlers behave, and how to reach us. What happens to the data after we fetch it is covered by our data use page.
Who we are
Largesso is an art-discovery platform. We index public collections from museums, galleries, and art platforms so that people can find works and follow them back to the institutions that hold them. Every work we display carries a credit line and a link to the source institution's own page for it.
How our crawlers identify themselves
Traffic from us arrives under one of these User-Agent strings:
largesso-pipeline/1.0 for catalog metadata, and
Mozilla/5.0 (compatible; LargessoBot/1.0; +https://largesso.com) for image
fetching. Both carry a contact address. We do not disguise our crawlers as browsers, and
traffic from us never arrives under a spoofed identity.
How they behave
Where an institution publishes an API, bulk export, or data dump, we use it instead of crawling pages. Where an institution documents a rate limit, that limit is written into our fetcher as a hard per-host floor it cannot cross. Hosts without a documented limit get a conservative default rate and a small connection cap.
If a server blocks us, we stop and ask. We do not rotate proxies or otherwise disguise traffic to get around a block, and a source that has asked us to stop stays off the platform until it gives explicit permission.
What we fetch
Catalog metadata (titles, artists, dates, credit lines) and preview images at modest resolution. How that material is hosted, embedded, attributed, and taken down is described on the data use page, and the per-source state is public at largesso.com/sources.
Contact and opt-out
Questions about our crawling, a rate you would prefer we use, or a request to stop entirely: email [email protected] or reach us through support. One email from the institution removes the source the same day; no legal letter required.