Crawler
VistasonRegistrySync
If you have found this address in your access log, this is the page it points at. VistasonRegistrySync is a crawler operated by Vistason. It reads the public websites of aviation companies in our directory — repair stations, distributors, operators and manufacturers — to record what each one publishes about its own capabilities. It reads a small number of pages, slowly, and it obeys robots.txt.
Who runs it
Vistason, a marketplace for airline surplus material. The crawler identifies itself on every request as:
VistasonRegistrySync/1.0 (+https://vistason.com/registry-sync)
Questions, complaints and removal requests go to sourcing@vistason.com. That address is read by a person rather than a rota, so please allow a few working days for a reply. If you want the crawling to stop before then, the robots.txt rule below takes effect on our next visit without waiting for us to answer.
What it reads
The public pages of a company’s own website — the sort of page a customer would be shown: capabilities, approvals, certificates, locations and contact details. It reads nothing behind a login and submits no forms.
robots.txtis the first thing we read on any host, before any page. It is fetched once per host for the visit and governs everything that follows — documents as well as pages. A PDF under aDisallowis not downloaded. Ifrobots.txttimes out or returns a server error, we leave that host alone for the rest of the run: “I could not ask” is not “yes”.- Then one page: your homepage. We fetch it once to confirm the address really is this company’s, because an address on file is sometimes a directory profile somebody pasted in. If it does not answer we try the
www.form once and stop. - A group naming us ends the visit there. No page, document, address or contact from the site is read or kept. What we record is the address itself and the fact that we were turned away.
- At most 12 pages per site, chosen from the site’s own navigation, plus at most four documents linked from them — a capability list, a certificate, a quality manual.
- At least two seconds between requests to the same host, and longer if your
Crawl-delayasks for it. We honour a delay of up to sixty seconds; past that we keep waiting the sixty and read fewer pages, rather than either ignoring the request or spending an hour on one site. - Ten minutes per company, then it stops with whatever it has.
Pages and documents follow different rules, and the difference is worth stating. Page crawling stays on the company’s own registrable domain: a link that leaves it is not followed. Documents are the exception. A capability list or a certificate is often served from somewhere else — a content network, a shared file host — and we do fetch those. Before we do, we fetch that host’s robots.txt and honour it exactly as we honour yours, so a file host that excludes us excludes us for every company at once.
A single request is abandoned after fifteen seconds; a page is capped at 4 MB and a document at 25 MB. The crawler runs at a shared ceiling of ten requests a minute across every site at once, so a visit is a handful of requests spread over a few minutes.
How often we come back
A company we managed to read is not read again for ninety days. A visit that produced no answer about the company — the host did not respond, or no site could be found — comes round again in a fortnight instead. A site that turned us away in robots.txt is neither of those: a refusal is recorded as its own answer and the site is left alone for the full ninety days, after which robots.txt is read once more and, if the rule still stands, nothing else is fetched. An email to sourcing@vistason.com takes the domain off the list altogether.
What we keep
Against the company’s record in our directory: the website address; the capability lists, certificates and quality manuals it publishes, both as links and as copies of the files, held privately; the part numbers and repair capabilities those documents name; business contact details published on the site, such as a sales or quotations mailbox and the people the site itself lists with their published address; social links, a telephone number, the year the business says it was founded, the staff size it states, and a short description of what it does. Those company fields are filled in only where our record is empty — anything an operator or the company itself has entered is never overwritten, and neither is data from the aviation registers.
We also keep one verdict per company: whether it publishes a capability list, and where. That is the point of the exercise. A company that publishes none is recorded as such.
What the crawler does not do: it sets no marketing consent or lawful-basis field on any contact, it builds no profile of an individual, it adds nobody to a mailing list, and it neither reads nor stores anything about the visitors to your site.
How to opt out
Add a group for us to robots.txt at the root of each host you want left alone:
User-agent: VistasonRegistrySync Disallow: /
We match the product token VistasonRegistrySync case-insensitively, as RFC 9309 requires, and a group naming us takes precedence over User-agent: *. To exclude part of a site rather than all of it, list those paths under Disallow as usual.
You can also write to sourcing@vistason.com naming the domain. We will stop crawling it and delete what the crawler collected from it. If you would rather the company was not in the directory at all, say so in the same message and we will remove the record — though the approval and certificate entries also come from the public aviation registers, which we do not control.
Legal basis
We process this information under legitimate interests — keeping an accurate business directory of aviation companies from what those companies publish about themselves — and no consent is inferred from a page being public.
Where a published business contact names an individual, that person’s rights apply as normal, and a message to sourcing@vistason.com is enough to exercise them.
Our privacy policy lists public websites among the sources we draw on, but describes them as a way of checking what a customer has told us; it does not yet carry a line for building the directory itself, and we are adding one. Until it does, this page is the fuller account and it is the one we hold ourselves to. The policy is also still a draft, and says so at the top.