Legal
Web Scraping Policy
Hudson Labs Inc.
Last revised: September 4, 2026
Hudson Labs limits web scraping to public data sources and the public parts of websites, identifies itself where required, respects site terms and access controls, and never swarms or attempts to bypass CAPTCHAs, firewalls, paywalls, or logins. We do not scrape to recreate a dataset or a complete website.
1. Purpose and scope
This policy sets out how Hudson Labs approaches web scraping (also referred to as web crawling or automated data collection). It applies to all Hudson Labs personnel and to any automated collection of data from public websites and public data sources carried out by or on behalf of Hudson Labs — whether in the ordinary course of operating our platform or as part of a one-off Deep Dive research project. It does not apply to data that Hudson Labs licenses from third-party providers or receives through provider APIs or SFTP, which is governed by the applicable provider agreements.
2. Definitions
- Web scraping. The automated retrieval of information from websites or online data sources, including crawling, parsing, and programmatic collection.
- Public data source / public website. A source, or a portion of a website, that is accessible to the general public without a private login, paywall, or other access control.
- Limited scraping. Low-volume, narrowly targeted retrieval of specific public data points (for example, a small set of company reference fields), as distinct from systematic or bulk collection.
- More-than-limited scraping. Any collection beyond the limited category — broader in scope, higher in volume, or systematic — which triggers a detailed terms-of-service review before proceeding.
- Deep Dive. A one-off, project-specific research exercise that may involve broad public web searches and targeted retrieval (by scraping or public API) from miscellaneous public sources. Deep Dives are not conducted on an ongoing basis.
3. Core principles
Hudson Labs conducts web scraping only within the following limits:
- We scrape only public data sources and the public parts of websites. We do not scrape non-public data.
- We do not scrape behind private logins or paywalls.
- We identify ourselves where required, and we do not disguise or misrepresent our identity to gain access.
- We never swarm a site, and we do not attempt to override, bypass, evade, or defeat CAPTCHAs, firewalls, or other access controls.
- We never scrape with the intent to recreate a source's dataset or to replicate a complete website.
- Where we perform more than limited scraping, we review the source's terms of service in detail and confirm compliance before proceeding. If we cannot confirm compliance, we do not scrape the source.
4. What we scrape
Normal platform use
- Company reference details from SEC.gov — for example, sector, subsector, ticker, company identifiers, and company changes.
- Investor relations (IR) websites — to identify new events and disclosures and incorporate them into our systems where they are otherwise missing.
Deep Dives only (one-off, not ongoing)
- Broad-based public web searches and targeted retrieval — by scraping or via public APIs — of subsets of data from miscellaneous public sources. Examples include JASDEC (Japan Securities Depository Center), FRED (Federal Reserve Economic Data), the Federal Register, corporate websites, and public news sources. Compliance for each Deep Dive is evaluated at the time of the project.
What we do not scrape
- Any content behind a private login, account, or paywall.
- Any non-public area of a website.
- Anything collected for the purpose of reconstructing a source's dataset or a complete website.
- Sources whose terms do not permit our intended activity (see Section 5).
5. Compliance and legal review
Hudson Labs applies a tiered review:
- Limited scraping in the ordinary course. We operate under the standing principles in this policy and confirm, before collecting from a source, that the data is public and that our activity is consistent with the site's terms and any published access requirements (such as rate limits and identification requirements).
- More-than-limited scraping. We conduct a detailed review of the source's terms of service or use and confirm compliance before proceeding. Where compliance cannot be confirmed, we do not scrape.
- Deep Dives. Compliance is evaluated on a per-project basis at the time of the Deep Dive, because these are one-off and source-specific.
Worked examples
- SEC.gov — reviewed; we scrape in strict compliance with SEC.gov's terms and published access requirements.
- SEDAR / SEDAR+ — reviewed; we have determined that scraping is not compliant with its terms, and therefore we do not scrape it.
6. Operational safeguards
- Volume and rate. Our scraping is targeted and low-volume. We rate-limit and throttle requests, honor published rate limits and robots directives where applicable, and avoid patterns that could burden a site. If we detect that our activity could overwhelm a target, we back off or pause. We never swarm.
- Identification and traceability. We identify ourselves where required — for example, through a descriptive User-Agent and, where appropriate, contact information — so that a site operator can attribute requests to Hudson Labs and reach us if needed. We do not spoof User-Agents or otherwise conceal our identity. We identify using our User-Agent string and scraping contact address, e.g., info@hudson-labs.com.
- Access controls. Encountering a CAPTCHA, firewall, login, paywall, or other barrier is treated as a signal that automated access is not intended. We stop; we do not attempt to bypass or defeat the control.
7. Frequently Asked Questions
Q. Do you or your data sources engage in web scraping (web crawling or similar processes) to gather information or data?
A. Our licensed platform data sources do not consist of scraped data — core datasets are obtained from third-party providers under license and delivered via API or SFTP. Separately, Hudson Labs does perform limited, controlled web scraping of public sources: in the ordinary course, to enrich and maintain certain records (see Section 4), and, on a one-off basis, during Deep Dives.
Q. Do you conduct a compliance or legal review related to web scraping?
A. Yes. We apply a tiered review (Section 5): standing principles for limited scraping of public data; a detailed terms-of-service review and confirmation of compliance before any more-than-limited scraping. If we cannot confirm compliance, we do not scrape. For example, we scrape SEC.gov in strict compliance with its terms, and we do not scrape SEDAR/SEDAR+ because doing so would not be compliant.
Q. What do you scrape?
A. In normal platform use: company reference details from SEC.gov (e.g., sector, subsector, ticker, company identifiers, company changes), and investor relations websites to detect new events and add ones missing from our systems. In Deep Dives only (one-off): subsets of data from miscellaneous public sources, via scraping or public APIs — e.g., JASDEC, FRED, the Federal Register, corporate websites, and public news sources. We do not scrape to reconstruct a dataset or a complete website.
Q. Is your web scraping confined to the public parts of websites?
A. Yes. We only scrape public data sources and the public portions of websites.
Q. Do you web scrape behind a paywall or login?
A. No. We do not scrape any content behind a private login or paywall.
Q. Do you take any affirmative actions when web scraping (solving CAPTCHA, clicking “I agree” to ToUs, etc.)?
A. We do not solve or bypass CAPTCHAs, and we do not click through or circumvent access walls to reach gated content. We do not accept terms solely to defeat an access barrier.
- If agreeing to ToUs, are ToUs complied with? Where terms of use govern a public source we scrape, we review them and ensure our activity complies (Section 5).
- If solving CAPTCHA, how (e.g., clicking “I am not a robot” or solving a picture test)? We do not solve CAPTCHAs — neither “I am not a robot” checkboxes nor image / picture tests. A CAPTCHA is treated as a signal that automated access is not intended, and we stop.
Q. What is your practice if you encounter a firewall?
A. We treat a firewall (or other block or access wall) as a signal that automated access is not permitted. We do not attempt to bypass, evade, or override it; we stop scraping that resource.
Q. Do you monitor scraping activity to avoid overwhelming the target site?
A. Yes. Our scraping is targeted and low-volume; we rate-limit and throttle requests, honor published rate limits and robots directives, and back off if our activity could burden a site. We never swarm.
Q. Are the methods used for scraping traceable back to you such that the target site could contact you (directly or via a proxy server provider) if needed?
A. Yes. We identify ourselves where required (e.g., via a descriptive User-Agent and, where appropriate, contact information) and do not disguise our identity. A site operator can attribute requests to Hudson Labs and reach us; to the extent any intermediary or proxy is used, activity remains traceable through that provider.
8. Governance and review
- Ownership. This policy is owned by the Hudson Labs Privacy and Security Officer, who is responsible for its application, for approving more-than-limited scraping and Deep Dive compliance evaluations, and for maintaining supporting documentation.
- Review. This policy is reviewed at least annually and updated as our practices or a source's terms change.
- Contact. Questions about this policy or about Hudson Labs' scraping activity may be directed to info@hudson-labs.com.