Back to the product·Privacy·Terms·Data processing

This is a draft, and no lawyer has read it. It was written from the code rather than from a template, so what it describes is what the system does — but the entity behind the service, the governing law, the notice address and the supervisory authority are still blank, and the mailboxes it names have to exist before it goes live. Each document ends with its own list of what is open. Do not treat any of this as executed terms.

Privacy

Draft of 14 August 2026 · read against the code on the same day

Two kinds of people are in this system and only one of them chose to be. Customers opened the product and used it. Everybody else is a person a customer researched, or wrote to, who never asked us for anything. The second group comes first here, because they have the least reason to read the rest.

If we wrote to you, or you found yourself in somebody’s list

What we hold about you is one row. Your name, the company you work for, an address at that company’s domain, a link to the page or source it came from, and one word for how it was established:

  • published — your employer put it on the web, and the row names the page.
  • smtp-confirmed — the mail server that runs your company’s mail answered that the mailbox exists.
  • inferred — we worked out the shape of your company’s addresses and rendered your name into it. Nobody has confirmed it. It is a guess, and the row says so rather than showing a green tick.

We do not hold your home address. We do not hold a phone number unless a business directory published it as the company’s. We hold nothing about what you have bought, read or browsed, no behavioural profile, and no advertising identifier — you were never on this website, so there is nothing to have collected.

To stop the mail

Two ways. Neither needs an account, and they do not work at the same speed.

  • The link at the bottom of the message. It opens one page with one button, and pressing it records the request there and then. The button is a POST, so a corporate link scanner that opens every URL in an inbound message cannot opt you out by accident — and cannot fail to opt you out either, because the button is the whole page. If the page says it could not record it, it means that, and the reply route below is the fallback.
  • Or reply and say so. A reply containing unsubscribe, remove me, take me off, stop emailing me, do not contact me or their Russian equivalents suppresses your address. It is matched on the words, by a list anybody can read, not by a model. Our own footer is stripped out before the match, so quoting the message back does not count as us asking ourselves. It takes effect only once the sending mailbox is actually read, and nothing reads it on a schedule today — the poll exists, and a person has to start it. Treat a reply as a request that will be honoured late rather than at once, and use the link if you have one.

Either route writes your address to a suppression list the sender checks before every message, including a reply inside a thread you had already opened. It is checked before the daily cap and before a mailbox is even chosen, and at that point it has no exceptions.

What it covers: everything from whoever wrote to you. The list is keyed by a scope as well as by an address. An opt-out from the link or from a reply is written against the account that sent you the message, so it refuses every campaign and every mailbox that account has — not only the one you were written from. It does not reach other customers of this service: each of them holds their own list, and one customer cannot be shown, or be refused by, another’s. That is a deliberate limit and it cuts both ways, so it is stated rather than glossed: if a different customer of ours later writes to you, you will need to stop them the same way. There is a service-wide scope and we use it only for addresses we must never write to at all. If you would rather have your whole company covered than your own address, write to the address in the next section and say so; the list holds a domain as readily as a person.

And no customer can undo it. There is a way to take an address off a do-not-contact list, because the word-matching above is deliberately generous and does sometimes read an ordinary reply as an opt-out. It needs a person — no agent in the system can do it — it requires a written reason, and it is recorded with who did it and when. It reaches only the entries a customer put on their own list. An entry that came from your click, or from your reply, is not one of those, and that removal cannot touch it.

To have the row deleted

Write to privacy@skygen.ai from the address, or naming it, and we delete it. Two limits, stated because they are true rather than because they are good:

  • There is no self-service delete and no automated erasure job. A person does it by hand. That is on the list of what to build; today it is a mailbox and somebody reading it.
  • Your suppression entry survives the deletion, for the reason above. It holds your address and the date you asked, and nothing else.

To be told what we hold

Ask, at the same address. You get the row itself, the URL or source it came from, the word above that says how it was established, and — if a message went out — when, from which mailbox, and the text.

To keep the crawler off your site

Disallow SkygenResearchBot in your robots.txt. It is fetched once an hour per host and obeyed, including Crawl-delay, which most crawlers ignore. Or write to bot@skygen.ai — the address is on every request we make, for exactly this.

To object, or to complain

You can object to the processing described here at any time, and we stop — that is what the suppression list and the deletion above are. You can also complain to your national data protection authority. If you are in the EU or the UK, the rights you are exercising are those in Articles 15, 17, 18, 20 and 21 GDPR.

Where a row comes from

Nothing here is bought in bulk and nothing is assembled from a logged-in scrape of a professional network. Each source below is asked about a company or a person that a customer named.

SourceWhat it gives
The company’s own websiteAddresses and names it published — contact, about, team, staff pages. Read by a crawler that fetches robots.txt first, obeys it, waits the delay the site asks for between requests — at least 0.4 seconds, and up to ten — identifies itself, and caps how many pages it takes.
Web searchNames, from the titles of results. Profiles are never fetched: only the search engine’s own result list is read. Through SearXNG when a deployment runs its own, otherwise Tavily.
OpenStreetMap (Overpass, Nominatim)Business name, location, and the phone or email the business published in the map itself.
Applicant tracking boards, Hacker NewsCompanies that are hiring, and what for. A signal with a date on it, not a person.
National company registers (FR, NO, DK, GB)Directors and owners, which those states publish by law. Matched on the company name before anything is attributed, because a wrong director is a confident sourced falsehood.
Public GitHub commitsThe name and address an author wrote into a commit they published.
WordPress author endpoints, German Impressum pagesNames a site publishes about itself through its own software or by legal requirement.
YouTube Data API, public TikTok profilesCounters for a creator account — followers, views, post dates. Numbers about an account, not contact details about a person.
Paid finders, only when configuredHunter, Prospeo, Findymail, AnymailFinder, LeadMagic. Asked about one named domain or one named person. No bulk list is bought, and none is uploaded.
A model reading one page, only when configuredScrapeGraphAI, last in the chain, on a page of the company’s own site. What it returns passes the same checks as anything else: the address has to be on the page it cites and at the company’s own domain.

There is one quarantined exception to the politeness rules above. A session lane reaches Instagram and X with a logged-in session, on its own machine and its own address, and it is off unless a deployment configures it. Everything it returns is marked scraped rather than observed, so a number only our session ever saw is never confused with one anybody can re-check by opening the page.

What an address check actually does

When verification is configured, we open an SMTP connection to the mail server your domain nominates and have the first half of the conversation a sender would have: EHLO, MAIL FROM, RCPT TO with the address in question, then QUIT. The server answers whether that mailbox exists. Nothing is delivered. The conversation stops before DATA, which is the line between asking whether a mailbox exists and putting something in it.

Three rules bound it, and they are in the code rather than in this paragraph:

  • Twenty recipients per domain per hour. Past that, the answer is unknown. The shape of a dictionary attack is many recipients at one server in a short time, and a probe budget is what stops us resembling one.
  • A control probe first. An address that cannot belong to anyone is tried before yours. If the server accepts that too, the domain accepts everything, and every other answer from it is reported as unproven rather than as verified.
  • A refusal to answer is never a verdict about you. Greylisting, a rate limit, a block on our address — all of those are facts about us. They are recorded as unknown, never as invalid, because invalid would delete a real person from somebody’s list on the strength of our own outage.

What this system will not do

  • The crawler never disguises itself. No spoofed browser fingerprint, no rotated proxies to get around a block, no CAPTCHA solving. It sends a name and a mailbox to complain to, and stops when told to. That is how the code is built rather than a policy that could be quietly relaxed.
  • A compliance mailbox is never a send target. gdpr@, dpo@, datenschutz@, privacy@, abuse@, postmaster@, security@, whistleblower@ and their translations are recognised and refused — including when they arrive bundled with another address in one field, which is how the rule was defeated once and then fixed. Every route that promotes an address to a send target passes the same gate.
  • Messages carry no tracking pixel and no rewritten links. Opens are not measured and we do not ask the mailbox vendor to measure them. What goes out is the text plus the opt-out line.
  • Nothing is sold, rented or shared with an advertising network. There is no advertising in this product and no analytics script on this site.
  • No special-category data — health, beliefs, politics, union membership, sexuality — is collected, inferred or stored. Nothing here looks for it.

The legal basis, and what it does not cover

For business contact data we rely on legitimate interest, Article 6(1)(f) GDPR: reaching a person in a professional role, at their employer’s domain, about their work. The balancing test rests on things that are true of the code — every message carries an opt-out, appended by the sending layer rather than by whoever wrote the message, so it cannot be left off; a person is written to once; a reply ends the sequence; compliance desks are excluded outright; the verification probe never delivers anything; the source of every address is recorded so anyone who asks can be told where it came from; and volume is capped low enough that no recipient sees a burst.

Two of those depend on something that is not yet automatic. A reply ending the sequence, and a reply that says stop being honoured, both require the sending mailbox to be read, and that read is not on a schedule — see the note under “to stop the mail” above. The link in the footer does not depend on it. We would rather say this here than let a balancing test lean on a step a person has to remember to run.

What that basis does not cover, said plainly:

  • Consumer marketing. This is not built for it and the terms forbid it.
  • Whether a particular message may lawfully be sent. Legitimate interest is our basis for holding business contact data. Whether an unsolicited commercial message may be sent to a given person is a separate question, and the answer differs by country — Germany, Austria and Italy among others require prior consent for much of what other countries allow. The customer chooses the audience and presses send, and under the terms that decision is theirs. Not yet: the product does not check a recipient’s country against that country’s rule, and nothing in the interface warns about it.
  • Scoring or profiling a person. Creator numbers are counters about an account. Nothing infers traits about an individual.

If you are a customer: what the product keeps about you

  • Your account. Email address, the date it was made, and a password hash — PBKDF2-HMAC-SHA256, a fresh salt per record, and the iteration count stored beside the hash so it can be raised later without locking anybody out. The password itself is never stored and cannot be recovered from what is.
  • One cookie. skygen_sid: 256 random bits, HttpOnly, SameSite=Lax, Secure off localhost, thirty days by default. It carries no claims — it is an unguessable number that says which browser is asking. There is no analytics cookie, no pixel and no third-party script on this site.
  • What your browser keeps. The engine also writes its own state to your browser’s local storage, so a refresh does not lose what you built. That copy never leaves your machine and clearing site data removes it.
  • Your engine state. The worlds you created, the commands you ran, the approvals you gave or refused.
  • Your memory. What you asked the engine, the tool calls each request actually produced, the companies and people you are working, notes per channel, and the standing facts you told it about yourself — your offer, who you sell to, your tone, domains you never want contacted.
  • What was sent. Recipient, which mailbox, which day, the thread, and a preview of any reply, up to five hundred characters. This is what enforces the cap, the one-person-once rule and the reply stop, so it cannot be discarded without those becoming decorative.

The corpus is shared between accounts

This is the disclosure that no template would make, and it is deliberate design rather than an accident.

Everything the engine works out about the world — that a company forms addresses as first.last, that a particular address was confirmed by a mail server, that a company exists and what it does, the people found at it, a creator’s public counters — goes into one shared corpus. It is not partitioned per customer. A convention learned once is true for everybody, and the whole commercial argument is that the answer gets cheaper the second time somebody asks. So a search by one customer can return a person that another customer’s run put there, and a row outlives the account that caused it.

What is not shared is everything about the account: your requests, the companies you are working, your notes, your standing facts, what you sent and to whom. That lives in a separate store, keyed per account, and nothing reads it without an account to scope it. Folding the two together would be a data leak wearing a performance optimisation.

For the corpus we are the controller in our own right, not a customer’s processor — which is why a customer cannot instruct us to delete from it, and why a person in it deals with us directly, at the addresses above. The data processing addendum says the same thing in the place where it matters contractually.

Who else sees any of it

Every row below except the first is optional and inert unless the deployment sets its key. With no keys at all the pipeline runs on free routes and reports which providers it would have asked rather than pretending it asked them.

WhoWhat reaches themWhen
CloudflareHosting, the D1 database, the shared store and the request logs. Everything stored passes through them.Always — this is where the product runs.
OpenRouter, and the model it routes toText from pages we fetched, the request you typed, the names and addresses found on a page, the letter being drafted — and, when a reply is classified, the words the person who replied wrote back, including their address and anything they put in that message.With OPENROUTER_API_KEY. Without it those steps return nothing and the planner says so in words.
TavilyThe search query: a company name, a domain, a person’s name.With TAVILY_API_KEY, and only when no SearXNG is configured.
SearXNGThe same queries.Self-hosted, so not a third party when you run the box.
OpenStreetMap (Overpass, Nominatim)The category and place asked about.Free routes, used by default.
GitHubThe organisation or repository being read.Always for public data; GITHUB_TOKEN only raises the rate limit.
Google (YouTube Data API)Channel and video identifiers.With YOUTUBE_API_KEY.
Hunter, Prospeo, Findymail, AnymailFinder, LeadMagicA domain, and a person’s first and last name.Each only when its own key is set.
ScrapeGraphAIOne URL on the company’s own site, and the question asked of it.With SCRAPEGRAPH_API_KEY.
InboxKitThe sending mailbox, the recipient, the subject and the whole message — and the replies, when they are read back.Whichever mailbox vendor the deployment configured. Without one, nothing can be sent at all.
The verification boxThe addresses being probed, and the domain they belong to.A machine the deployment runs. Without it nothing is verified, and the product says that rather than guessing.
The session-lane boxAn Instagram or X handle.Off unless configured, and never on the machine the polite crawler uses.

Where it lives, and for how long

Data is stored by Cloudflare — D1 for engine state and the outbox, their shared store for the corpus, per-account memory and account records. We have made no data-residency commitment and do not offer one yet.

WhatHow long
Session cookie and its recordThe cookie is issued to expire in thirty days, after which your browser stops sending it. If you signed in there is also a record on our side, and signing out blanks it, which is what revokes the session. If you never signed in there is no record on our side at all — the cookie is the whole of it, and the thirty days is your browser’s doing rather than ours.
Worlds — engine state, commands, approvalsMeant to be swept hourly, deleting anything untouched for thirty days. That sweep is a timer inside a long-running server process, and the edge runtime this is deployed to does not keep one — nothing schedules it there today. So dormant worlds are not being collected on any schedule we can stand behind, and deletion on request is the route that works. Written here rather than left as an assumption, because a retention period nothing enforces is not a retention period.
Account record and per-account memoryUntil you ask us to delete it. There is no self-service delete yet.
The shared corpusNot swept. Confidence halves at a fixed rate — every 180 days for a company’s address convention, every 90 for a confirmed address, every 90 for a published person, every 45 for an inferred one — and a row stops being used and stops being returned once it has fallen below the floor the caller asked for. Those are half-lives, not expiry dates: a row is faded rather than gone on the day the number lands. Decay is not deletion. The row stays in the store until somebody removes it, and the way to have that done is the mailbox above.
Creator readingsUp to 730 per creator — two years of daily readings — kept on purpose, because a growth curve cannot be backfilled.
Sends, replies and reply previewsKept while the account exists. They are what enforce the send-once and reply-stop rules.
Suppression listIndefinitely, deliberately. It is the record that somebody asked not to be written to.

Logs

The service writes one structured JSON line per request to standard output, captured by Cloudflare. Most lines carry a domain, a status and a duration. Some carry an email address — a suppression is logged with the address it suppressed, which is the point of that line. Not yet: addresses in log lines are not redacted, we run no log store of our own, and we have set no retention beyond the hosting platform’s default.

Not for consumers, and not for children

This is a business tool that contacts people at work. It is not offered to anyone under 16, it is not for marketing to consumers, and it must not be used to research or reach a private individual as an individual — which is a rule in the terms as well as a sentence here.

Changes

When this changes, the date at the top changes and the change is described here. A change that widens what is collected or who receives it is announced to customers before it takes effect, not after.

Contact

privacy@skygen.ai for anything on this page — access, deletion, objection, or a question about a row. bot@skygen.ai to tell the crawler to stop. security@skygen.ai for a vulnerability.

What is still open in this draft

  • The legal entity, its registered address, and the supervisory authority it answers to.
  • A named representative in the EU or the UK, if one is required.
  • A scheduled read of the sending mailboxes, so “reply with the word unsubscribe” is honoured without a person starting it, and so a reply reliably ends a sequence.
  • A collector that actually runs for worlds nobody has touched in thirty days, so the retention table above describes something enforced rather than intended.
  • An automated erasure path. Today deletion is a mailbox and a person.
  • Redaction of addresses in log lines, and a stated log retention.
  • A data-residency commitment.
  • A check of the recipient’s country against that country’s rule on unsolicited commercial mail.
  • Confirmation that the mailboxes named above exist and are read.

Our crawler identifies itself as SkygenResearchBot/1.0 and sends From: bot@skygen.ai on every request. To keep it off a site, disallow that name in robots.txt or write to that address; both are honoured.