DPAFlow crawlers
We monitor publicly published subprocessor lists so companies can meet their GDPR Article 28 obligations. If you run one of those sites, this page tells you what we send, where it comes from, and how to control it.
What we send
Covers the request classes we have identified as reaching a site we monitor or verify. It excludes traffic to our own service providers and to customer-configured webhooks. If you see something from us that is not described here, tell us and we will add it. 7 of the 13 classes below carry a DPAFlow token you can match on; the other 6 do not, so for those the addresses further down are how you recognise us.
| Request class | User-Agent | Purpose |
|---|---|---|
| DPAFlow Source ScanSource-scan workerIdentifies as DPAFlow | Mozilla/5.0 (compatible; DPAFlowSourceScan/1.0; +https://dpaflow.com/bot) | Reads the public subprocessor list page a customer monitors, including downloading it when it is a PDF. |
| DPAFlow ProbeBrowser-probe worker, first attemptIdentifies as DPAFlow | Mozilla/5.0 (compatible; DPAFlowProbe/1.0; +https://dpaflow.com/bot) | Checks that a candidate URL is reachable and is the page it claims to be, before we monitor it. |
| DPAFlow Probe (rendered)Browser-probe worker, second attempt (headless Chromium)No DPAFlow token in the User-Agent | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/<version> Safari/537.36Varies with the browser build — match on the shape, not an exact version. | Retries the same check in a real browser when the plain read did not settle it — which is often straight after a site blocked the first attempt. Carries no DPAFlow token. |
| DPAFlow Source VerificationScheduler, and super-admin remediation actionsIdentifies as DPAFlow | DPAFlowSourceVerification/1.0 | Re-checks a source URL after it was reported broken, moved or blocked. |
| DPAFlow Website VerificationDashboard, when a customer adds a websiteIdentifies as DPAFlow | DPAFlow-WebsiteVerification/1.0 (+https://dpaflow.com) | Confirms that a website a customer registered is reachable and belongs to them. |
| DPAFlow Website Technology ScanDashboard, or the scheduler on a scheduleIdentifies as DPAFlow | DPAFlow-WebsiteTechnologyScan/1.0 (+https://dpaflow.com) | Reads a customer's own registered website to record the technology it uses. |
| DPAFlow Website Legal-Page ScanOnboarding website scan, through the shared page fetcherIdentifies as DPAFlow | DPAFlow-VendorTechDetector/1.0 (+https://dpaflow.com) | When a customer asks us to scan their own website, reads its home page, follows the links on it that look like legal or privacy pages, and reads a small number of those. Shares the technology detector's User-Agent. |
| DPAFlow URL discoveryURL-discovery workerNo DPAFlow token in the User-Agent | nodeNot set by us — this is our HTTP client’s default and can change when we upgrade our runtime. | Looks for a vendor's subprocessor page by reading robots.txt and sitemap.xml and trying a short list of common paths. Sets no User-Agent of its own, so our HTTP client's default goes out instead. |
| DPAFlow evidence captureProtected evidence-capture workerNo DPAFlow token in the User-Agent | nodeNot set by us — this is our HTTP client’s default and can change when we upgrade our runtime. | Re-reads an approved source page to store the evidence snapshot. Sets no User-Agent of its own, so our HTTP client's default goes out instead. |
| DPAFlow screenshot captureScreenshot worker and source-scan evidence capture (headless Chromium)No DPAFlow token in the User-Agent | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/<version> Safari/537.36Varies with the browser build — match on the shape, not an exact version. | Renders the same public page to store a dated screenshot as evidence that the disclosure existed. |
| DPAFlow browser validationHermes agent, including its proof-screenshot capture (headless Chromium)No DPAFlow token in the User-Agent | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/<version> Safari/537.36Varies with the browser build — match on the shape, not an exact version. | Loads a public page in a real browser when a plain HTTP read cannot render it, and captures the proof screenshot for it. On this path a challenge page is recorded and the read stops. |
| DPAFlow portal loginHermes research agent, when an operator configures portal credentialsNo DPAFlow token in the User-Agent | Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:128.0) Gecko/20100101 Firefox/128.0Varies with the browser build — match on the shape, not an exact version. | Signs in to a vendor portal to read documents that are not published publicly. The account is one DPAFlow operates and an operator configures, not one the site owner set up for us, and the sign-in flow can register an account where none exists. It runs only for a job an operator authorised for account login, against a domain on the operator's allow-list, and it presents a desktop Windows User-Agent rather than the browser it really runs; see the page. |
| DPAFlow URL verification toolOperator tooling, run by handIdentifies as DPAFlow | Mozilla/5.0 (compatible; DPAFlowBot/1.0; +https://dpaflow.com/bot) | Checks a batch of candidate source URLs when we are curating the vendor catalogue. It runs from whatever machine the operator uses, so it has no fixed address. |
Where we come from
92.4.90.125source-scan, browser-probe, browser-probe-render, website-technology-scan, url-discovery, evidence-capture, screenshot-capture, browser-rescue178.105.142.154website-verification, website-technology-scan, website-legal-page-scan
3 classes are not in that list — DPAFlow Source Verification, DPAFlow portal login, DPAFlow URL verification tool. Those have no address we can promise you.
These addresses can change, and an operator-configured route could add another or replace one. Re-read this document rather than treating it as a permanent allow-list, and use the User-Agent tokens as well where you can. The same data is available as JSON at /bot/ips.json, cached for one hour.
What we cannot state yet
- Our research agent can fetch a vendor page over plain HTTP without a browser. It leaves from 92.4.90.125 like the rest of the relay traffic, but its User-Agent is set by a third-party agent runtime we have not measured, so it is not listed above. Match it by address until we can state it.
Why we visit
GDPR Article 28 requires data processors to disclose the subprocessors they use, and to give customers notice when that list changes. Our customers are those data controllers. We read the page you already publish for exactly this purpose and alert our customer when it changes. See our privacy statement and security overview.
Control it
Please read this before relying on robots.txt. We do not currently parse Disallow rules, so adding one will not stop us today. We are telling you that rather than letting you believe a rule is working when it is not. What does work:
- Email [email protected] and we will stop, slow down, or switch to a URL you prefer. If our volume is a problem, this is the lever — we would rather hear from you than be blocked.
- Block or rate-limit the addresses above at your edge. That covers every class except the ones listed as having no fixed address, so pair it with the User-Agent tokens where you can.
We do intend to honour Disallow. When we do, this page will say so and these are the tokens it will use:
User-agent: DPAFlowSourceScan
Disallow: /
User-agent: DPAFlowProbe
Disallow: /
User-agent: DPAFlowSourceVerification
Disallow: /
User-agent: DPAFlow-WebsiteVerification
Disallow: /
User-agent: DPAFlow-WebsiteTechnologyScan
Disallow: /
User-agent: DPAFlow-VendorTechDetector
Disallow: /
User-agent: DPAFlowBot
Disallow: /Challenges, and the one path that is different
On every request class listed above — the ones that read a page you publish — a CAPTCHA or WAF challenge ends the attempt. We record that we were challenged and stop. On those classes we do not try to get past it and we do not call a CAPTCHA-solving service.
There is one path where that is not the whole truth, and we would rather write it down than let the paragraph above stand as an absolute. Portal sign-in is a separate, operator-configured path that signs in to a vendor portal to read documents that are not published publicly. On those jobs our agent is permitted to work through a challenge that stands between it and the account it is signing in to, using a browser-automation skill whose capabilities we do not enumerate here. That path also presents a desktop Windows User-Agent rather than the browser it really runs, with the usual automation marker hidden.
Two things about that account, because they are the parts people assume otherwise: it is an account DPAFlow operates and an operator configures — not one you set up for us — and the sign-in flow can register an account where none exists. If you would rather we did not hold one on your portal, email us and we will remove it.
Where that path stops, because the boundary matters more than the exception: it runs only for a job an operator has authorised for account login, and only against a domain on the operator’s allow-list. A scan, a discovery run, a screenshot or an ordinary research job cannot reach it, whatever the rest of our configuration says. And nothing it meets on a challenge page can become evidence about you — a page that looks like a challenge or a login is rejected before it reaches the record we keep, not stored and labelled.
Last updated 26 July 2026.