Public Telegram posts, personal data, and who the controller is
A public post can contain an email, a phone number and a handle, and returning it in full is what makes it useful and what makes it personal data. Where that leaves the person running the run.
Public Telegram channel posts are published to the open web, and that does not stop them being personal data when they contain an email address, a phone number, an @mention or a t.me handle — which measured 3.0% of posts overall and 9.0% in job channels. Whoever runs the Actor is the data controller for the output and needs their own lawful basis under Article 6 GDPR if they are in scope. The Actor's own position is narrow by design: it returns what a channel published to whoever asked for that channel, does not cross-reference authors between channels, does not enrich from outside sources and does not build profiles.
Key points
Public availability is not a lawful basis. Article 6 GDPR requires one regardless of how the data was published.
Whoever runs the Actor is the controller for the output, because they choose the purpose and the means of the processing.
Contact entities are measurable rather than hypothetical: 3.0% of posts overall carried an email, 9.0% in job channels, and 35% in one job channel alone.
`includeContactFields` set to false empties the email, phone, mention and handle columns; the post text is still returned in full, because a post is what the channel published.
The Actor does not cross-reference authors between channels, does not enrich from outside sources, does not build profiles and does not discover channels by following mentions.
The most common mistake in this area is treating “it was already public” as the end of the analysis. Article 6 GDPR requires a lawful basis for processing personal data, and it does not contain an exemption for data the subject published themselves. Public availability settles whether you can read something; it does not settle what you may then do with it.
Who the controller is
Whoever runs the Actor. They choose which channels to read, for what purpose, and what happens to the rows afterwards — which is what determines the purpose and means of the processing. The Actor is a tool that executes that choice.
What is actually in a post
Not a hypothetical. Across 893 measured posts, 3.0% contained an email address — 0% in news channels, 9.0% in job channels, and 35% in one job channel measured on its own. @mentions were on 30.8% of posts and t.me handles on 43.2%, rising to 53.5% and 64.2% in job channels.
A job posting that says “send your CV to name@company.com, ask for @recruiter” is a public post and it is also personal data about at least two people.
The switch, and what it does not do
includeContactFields: false empties four columns: emails, phone numbers, @mentions and t.me handles. It does not redact the post text, and that is deliberate — a post with its contacts cut out of the body is no longer the post the channel published, and silently altering source text is worse than returning it whole and saying so.
Four things the Actor refuses to do
It does not cross-reference authors between channels. It does not enrich anything from outside sources. It does not build profiles of people. It does not discover channels by following mentions.
Each of those is a step from “what this channel published” towards “a file about a person”, and the line is drawn before the first of them rather than after the third.
Retention and removal
Cached pages are held for 7 days and archives for 90, both expiring automatically rather than by a cleanup job somebody has to remember to run. Every row declares from_cache, fetched_at and data_age_hours. Removal requests go to the address on the data removal page.
Frequently asked questions
▸Is scraping public Telegram channels legal?
The posts are published to the open web and readable without an account, which settles access rather than processing. If the output contains personal data and you are within the GDPR's scope, you need a lawful basis under Article 6 for what you do with it — public availability is not itself one. This is a developer's description, not legal advice.
▸Who is the data controller?
Whoever runs the Actor. They choose which channels to read, why, and what happens to the rows afterwards, which is what defines the purpose and means of the processing.
▸Can I exclude contact details?
Yes. Setting `includeContactFields` to false empties the emails, phone numbers, @mentions and t.me handles columns. The post text is still returned in full, because a post with its contacts stripped out of the body would no longer be the post the channel published.
▸How long is anything retained?
Cached pages for 7 days and archives for 90, both expiring automatically. Every row declares `from_cache`, `fetched_at` and `data_age_hours`. Removal requests go to the address on the data policy page.
Sources
Every URL below was requested and returned a page on the date shown.
A walkthrough of reading public Telegram channels anonymously: the four modes, when search beats reading a history, and what the anonymous preview does and does not render.
The Bot API cannot read a channel it is not in. MTProto needs a phone number and an api_id. The public web preview needs neither, and it is what a channel publishes to the open web.
Documents, voice notes, audio, stickers, locations and round videos never appeared once in the anonymous preview. Fields that would always be false were cut rather than shipped.