The principle
HuiTu Technology works with publicly available information, openly licensed data, and data that a client is authorised to use. That boundary is the basis of the whole practice, not a disclaimer attached to it.
It exists for two reasons. The first is that we do not want to do the other kind of work. The second is commercial: data obtained by circumventing protections stops working, often without warning, and the client inherits both the gap and the risk.
Sources we work with
- Publicly accessible web pages that do not require authentication to view.
- Open government, statistical and municipal datasets, under their published licences.
- Open collaborative mapping data, with the attribution its licence requires.
- Official APIs, within the terms of the provider.
- Licensed commercial datasets, where the client holds or purchases the licence.
- Data the client owns or is otherwise authorised to provide to us.
Work we decline
We say no during scoping rather than after a contract is signed. Where we decline, we normally suggest a lawful alternative: an official API, a licensed provider, or a different method that answers the same question.
- Anything requiring the circumvention of authentication, paywalls, rate limiting or anti-bot protections.
- Collection from sources whose terms clearly prohibit the intended use, where no alternative route exists.
- Building datasets about private individuals, including profiling, tracking or aggregating personal information.
- Collection intended to harass, discriminate against, or misrepresent an identified person or group.
- Reconstructing a licensed commercial dataset in order to avoid paying for it.
How we collect
- We request at a rate that does not degrade service for anyone else, and we back off when a source signals load.
- We identify our traffic honestly and do not disguise it as something it is not.
- We collect the fields the project needs, not everything a page happens to contain.
- We store the minimum necessary and delete working copies when a project closes.
- We record what was collected, when and from where, so provenance is always traceable.
Personal data
Business contact details published by a business about itself are treated as business data. Information about identifiable private individuals is not collected unless it is strictly necessary, lawful for the stated purpose, and agreed explicitly with the client in advance.
We do not build consumer profiles, we do not collect or process device or mobility data that could identify individuals, and we will not accept a project whose purpose is to track people.
Attribution and licensing
Open data usually carries an attribution requirement, and some carries a share-alike condition that affects what you can do with a derived product. Where third-party data is included in a delivery, the source, its licence and the required attribution are stated in the delivery documentation.
It is worth reading that section before publishing anything built on the data. Attribution obligations are easy to satisfy and awkward to fix retrospectively.
Client-supplied data
- We use client data only for the project it was supplied for.
- It is never combined into products for other clients.
- It is stored on access-controlled systems for the duration of the engagement.
- It is deleted on request, and in any case once it is no longer needed.
Accuracy and honesty about limits
No collected dataset is complete and no analysis is free of assumptions. We state the achieved coverage, the validation performed, the known gaps and the assumptions relied on. We do not claim guaranteed accuracy or completeness, because those claims are not true of any dataset of this kind.
Questions and concerns
If you operate a website and have a question about traffic you believe originated with us, or if you have a concern about how data relating to you may have been used, contact jinlongc7c@gmail.com. We respond to these directly and promptly.