Your child's data profile doesn't start when they get their first phone. It starts before they're born, the moment a parent emails a gynecologist or visits a fertility clinic website. That's the core argument behind Born Private, Proton's new initiative that lets parents reserve an email address for their child at birth, anchoring their digital identity in a privacy-preserving ecosystem before the profiling machine gets started. Craig Smith sits down with Eamonn Maguire, Engineering Director, Machine Learning & AI at Proton, who has spent his career at the intersection of data, security, and visualization to explore what's really happening to our data and what, if anything, we can do about it.
The conversation covers the mechanics of how just three email sign-ups can allow Google to infer your age, politics, and religion; why OpenAI and Anthropic have shown "not much regard for the law" when it comes to training data and copyright; and why social media platforms are operating like unregulated gambling companies - engineering addiction with no structural incentive to stop. It's one of the most grounded, specific, and genuinely alarming conversations about digital privacy you'll hear, and it ends with a simple, actionable proposition: privacy should be a decision you make at birth, not a problem you try to solve after the damage is done.
Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.
[00:00:00] [SPEAKER_01] My wife is convinced that her iPhone listens to the conversations because we'll have a conversation at dinner and then she'll get served an ad.
[00:00:08] [SPEAKER_00] The Profiles of Your Child is being created even before they're even born. There's no real transparency over to say how exactly these profiles are created in the first place. You think your child is safe in their bedroom, but all these companies are basically changing the behavior of your child.
[00:00:23] [SPEAKER_01] Do you think that there will develop an ecosystem around maybe Proton will be the kernel around an ecosystem of privacy to counter the Googles and Metas of the world so that someone with the born private email address could then participate in social media or participate in search without having them having somebody build this elaborate detailed profile of them.
[00:00:50] [SPEAKER_01] We're going to talk about born private, but Proton has a bunch of privacy preserving products. I thought maybe you could start by introducing yourself. I know that you've been working in this space for a long time. I believe you have a PhD from Oxford and that you worked at CERN. Is that right? Am I wrong on that? Yeah.
[00:01:16] [SPEAKER_00] No, no, you're not wrong. So that's correct.
[00:01:17] [SPEAKER_01] So, yeah. Can you introduce yourself and give that background?
[00:01:25] [SPEAKER_00] My background is varied, I would say. I started most of my professional type of career was in bioinformatics. It was basically using computational techniques to analyze the genetic sequences or protein sequences and so on. And that very much got me into the academic mindset of things. And I went and did a master's in bioinformatics and then a PhD in computer science.
[00:01:55] [SPEAKER_00] Largely focused on also quite a bit of bioinformatics, but I also worked in security applications. A lot of stuff in insider threat, for example. I worked in things with digital humanities also like analyzing poetry for the way things are spoken and how you can visualize that type of information. So I sort of my major interest for a lot of my life was working on visualization, data visualization.
[00:02:23] [SPEAKER_00] So the intersection of computer graphics, mathematics, statistics, machine learning later, I would say, and then the computer science aspect of it. And visualization opened up a lot of different avenues, in fact. So I worked in biology, I worked in security, I worked in digital humanities. I ended up doing my postdoc at CERN.
[00:02:47] [SPEAKER_00] I spent a couple of years there, then I went and worked in finance for three years, taking a lot of what I learned from, I would say, the previous, my previous career cycles into finance also. So I did a lot of work applying techniques typically used in biology or bioinformatics, but doing it using those techniques in financial time series analysis, for example, or looking at different ways of correlating different signals together.
[00:03:17] [SPEAKER_00] And after that, I ended up working at Facebook, where I worked a lot and with the machine learning and data science for all external and internal threats. So everyone trying to get into Facebook's network, preventing that happening. And then anyone trying to exfiltrate data out of Facebook, from within Facebook, also stopping that as well. So more detection engineering. And then after that, I came to Proton.
[00:03:46] [SPEAKER_00] So I've been at Proton Life for six years.
[00:03:48] [SPEAKER_01] Okay. I got to ask the digital humanities is fascinating. You're creating visualizations of language patterns in poetry. Is that essentially what you're doing?
[00:04:03] [SPEAKER_00] So one project, in fact, was looking at how the tongue moves. There's basically different ways. There's a representation about how you create sounds, like from the back of the throat, from the front, from the tongue and so on. And basically we had a grid representing that. And then you were able to visualize the changes of this over time.
[00:04:23] [SPEAKER_00] So you could see how different poems were supposed to be read in a more smooth way versus those which are more harsh tongues, where you have big dynamic changes in positions, which manifested themselves then as being very large changes in tonality, for example. So it was kind of cool. I like everything, which is interesting.
[00:04:49] [SPEAKER_00] And yeah, we worked on lots of stuff looking at, which is relevant also today for computing side of things, like looking how documents change over time. You know, like if you look at things like the Bible, for example, it's many manifestations of copies over many, many hundreds or thousands of years that, you know, the things were happening. So when you look at things like the other things that were added, things were removed. There's different languages.
[00:05:19] [SPEAKER_00] So you've got different translations. And looking at how these are done also in other types of texts as well. So like the additions and deletions, which then if you think in computing terms also, it's the same. And like looking at how files change over time is a complex thing or how code changes over time. And using that in security is also interesting. So we ended up using a lot of those things in security visualization problems as well.
[00:05:46] [SPEAKER_01] Wow. That's fascinating. And how did Proton get started? And what was its original mission?
[00:05:54] [SPEAKER_00] Well, the story goes that Proton got started in the CERN cafeteria. Yeah. And so Andy was doing his PhD at CERN. I think many of the founders were also still at the time doing their PhD at CERN in particle physics. And it was the time of the Snowden revelations. And Andy is, of course, his Taiwanese American descent.
[00:06:23] [SPEAKER_00] So for him, what was happening in Taiwan, for example, the threat of China, for example, but then also in America where it's a country where it was purporting itself to be the free democratic sort of flag bearer in some way. And then to understand just how deep the surveillance operations were going, I think it scared a lot of people at the time.
[00:06:50] [SPEAKER_00] And I think that Andy decided at that point in time to create something from mail because in the day-to-day, especially in science, I would say, it doesn't seem very scientific, but a lot of the day-to-day operations go around sending emails. And, you know, all of this IP that is being sent around is suddenly globally available to players like Google or Yahoo, for example, or Microsoft.
[00:07:16] [SPEAKER_00] And, you know, a lot of your personal life is attached to these emails. So starting it at email seemed to be a logical step. And that's when they started in 2014. And it was crowdfunded. So there was a Kickstarter project that is, I think, around half a million dollars.
[00:07:36] [SPEAKER_00] And that went a long way to setting up the company, bitstrapping everything to make it, you know, have the servers and have the capacity to handle the users. And then since then, it's funded by our subscribers. We don't have investors or we don't have venture capital. It's basically just funded by the fact that we have paying users and they really believe in the product. And they believe in what we're trying to do.
[00:08:04] [SPEAKER_00] And they, by paying for those subscriptions, they're able to also help bring privacy to other users, which are not paying. Because of course, Proton is also available for free to users as well. So across all of our products.
[00:08:20] [SPEAKER_01] I didn't realize that. So the business model is there's a freemium offering and then you pay for premium services or something.
[00:08:32] [SPEAKER_00] Yeah, for instance, for VPN, if you're a free user, you get access to certain countries. You don't have very much choice. But if you're a paying user, you've got access to hundreds of countries. I think something it's a lot of different gateways you can connect to. And for mail, for example, you have more storage. You have the ability to have custom addresses or custom domains.
[00:09:00] [SPEAKER_00] For Lumo, for the product I control or I run, basically you've got increased limits, better models and so on by paying. So we believe in the need to be able to make privacy available to all. But we also need to run a business as well, which is actually paying for all the people who are working on these things and all the infrastructure it requires to run it. And, you know, AI infrastructure in particular is quite expensive.
[00:09:25] [SPEAKER_00] So it's difficult to run that on a free model. Unless, of course, you're getting into things like advertising and so on, which is the way that most players are demonetized.
[00:09:38] Yeah.
[00:09:39] [SPEAKER_01] And we're going to talk about born private, this new product or project. And that fascinates me, which is that the parents can register an email address for their children at birth and that follows them throughout their lives. Right. But before we get to that, what are the risks, particularly in using AI?
[00:10:08] [SPEAKER_01] I mean, before we got on, we were talking about proprietary models. And I'm going to start paying more attention, but my understanding was that proprietary models, most of them say that they don't use your data for training. I can't. I know I've seen that notice on some of them. What happens with your models?
[00:10:38] [SPEAKER_01] And can you talk about, is it LUMO? Is that the name of your model, which is open source, but how do you protect privacy? What makes that different from other models? And if you don't have persistent memory, doesn't that limit the capability of the model? Can you talk about that?
[00:11:07] [SPEAKER_00] Sure. So first, like most models, so if you go to, as a consumer, let's say, if you go on to many of the big platforms like ChatGPT or Cloud and so on, you'll, by default, typically you give them permission to train on your data. Yeah. That's, that's their, I mean, if you're a free user, basically they need to make something from your free users, right?
[00:11:36] [SPEAKER_00] Otherwise why give you access to the platform unless you're not paying for it. If you're a paid user, then it's a bit different. If you're enterprise user, typically they all have it in the document and the contract saying we do not use your data to train models. But then it's also a trust problem, right? Because they do have access to all of your data. And not, not only all of your, the, the conversations that are happening back and forth between the user and the model, but also all the context that is provided as well.
[00:12:04] [SPEAKER_00] Right. And the business practice of these, at least at this point in time for all these providers, the, the real valuable thing is data that hasn't seen before, a text that has not seen before. Right.
[00:12:19] [SPEAKER_00] Because this is what makes the model better, right? It's not, you know, you can have all these fancy derivations of algorithms or mixture of expert types of architectures and, and small changes to, to how things are being processed or attention is being given. But at the end, the only thing that really matters is compute. So how much you're wanting to throw at it, but also how much data you have. And for them getting access to data is the most important thing.
[00:12:48] [SPEAKER_00] So that's why opening up to enterprise and education platforms is particularly of interest to these platforms because they want the data that hasn't seen. And, you know, within a corporate infrastructure, you know, inside proton, for example, we have probably hundreds of thousands of documents sitting around that have never been exposed to the public web and education is the same.
[00:13:11] [SPEAKER_00] So if you have a partnership between, for example, chat GPT and Oxford University, what do chat GPT get in return, right? Because they give favorable return, favorable terms to, to Oxford. And Oxford, and Oxford are presumably also giving back some data to them. But even if they didn't do it directly, all the students are still uploading all of their files into open, open AI servers that they have access to.
[00:13:40] [SPEAKER_00] And then you have to trust that they're not going to use your data. And I think that's a difficult thing to trust whenever you've seen that they basically threw away all sorts of regard for copyright law when they decided just to start scanning lots of books.
[00:13:55] [SPEAKER_00] And you see that with Claude Anthropic as well, where there is this $1.5 billion lawsuit against them because they basically went along and bought lots of thousands and thousands of books, scanned each of the pages and threw away all the books. Right. So that there'd be no paper trail because there were no digital trace. But at that time, they were going to buy all these books just in the shop. I mean, at least they bought the books, I would say.
[00:14:22] [SPEAKER_00] But then what they did afterwards is quite bad. And then you also have the same thing with Meta, for example, where they were found to have used this big archive of pirated books as well in order to train their models. So I think trusting ChatGPT or OpenAI or Anthropic with copyright,
[00:14:48] [SPEAKER_00] they haven't really shown themselves to have much regard for the law, I would say, at this point. Not to say that others are much different, but there are providers of models such as the Allen Institute, for example. So AI2 developing the open models, which are called ALMO. There's the Swiss AI initiative also, which has developed the Peritus, which is all based on open data too.
[00:15:16] [SPEAKER_00] NVIDIA are also creating their Nemetron models, which are all based on open data as well. So there are people who are creating open models, but truly open, not just open weights. But also you can see where the data comes from. You can see the code. You can see the process behind it. And I think open models has kind of been, what do you call it?
[00:15:43] [SPEAKER_00] Like the term has been degraded somewhat. It's open washing where people say something is open source or whatever, something is open. But actually it doesn't really satisfy the requirements of being open because just one small facet of it is.
[00:16:00] [SPEAKER_01] Yeah. That's like Lama, not as Lama. Is that right?
[00:16:05] [SPEAKER_00] It's open weights. Open weights, but you don't know where the data comes from.
[00:16:10] [SPEAKER_01] That's right.
[00:16:10] [SPEAKER_00] I mean, in the end, no one wants people to know that because we know, for example, for GPT-2, which is a long time ago now, there is like some crazy stat on the Wired article from around that time, which stated that the total data, so 0.3% of the training data used for GPT-2 can be accounted for the entire English language section of Wikipedia.
[00:16:39] [SPEAKER_00] So only 0.3% of the training data is basically everything on Wikipedia in English language. The rest, script web pages, social media profiles, whatever else before everything got locked down by Reddit and X and so on.
[00:16:56] [SPEAKER_01] You know, I remember hearing and repeating in the early days of these pre-trained transformer models that they had read the entire English language internet. And I never knew whether that was really true. It sounded unlikely. But what's your view? Is that possible even?
[00:17:24] [SPEAKER_00] Yeah, I think so. I mean, with the databases that you have, you've basically got Open Crawl, for example, that you can go and use. You can figure out which pages are interesting. You have to filter out all the garbage, of course, because there's a lot of garbage there, too, and scam websites and so on. But, I mean, knowing which pages are English language is pretty easy, mostly from the domain name, I would say. Yeah.
[00:17:51] [SPEAKER_00] And then going off and having giant crawlers getting all that content is also not a big deal. You know, the most impressive thing at the time for companies like OpenAI, for example, where, you know, the whole transformer model thing was around for quite some time. But no one really thought it would actually give decent outputs. And the fact that they went off and created GPT-2 with, you know, there's a lot of investment behind it.
[00:18:19] [SPEAKER_00] There was still, like, hundreds of millions of investment. So spending the money to go off and build sheep crawlers to go and get the content from the internet was not a big deal. So, like, if you know that the data is the king and the data is the oil for your machine, you make sure that you go and secure the oil. And that's what all these companies did. And all the companies that do the best in terms of having the best models are the ones that manage to acquire the most data.
[00:18:48] [SPEAKER_01] Yeah. I mean, it's so fascinating. I want to get back to Born and Private, but let me ask one more question. There's so much, you know, I do a lot of research, and there's so much that's not on the end of this, that's in on paper, that's not digitized.
[00:19:12] [SPEAKER_01] Do you think at some point all paper archives all over the world will be digitized and then available for training?
[00:19:21] [SPEAKER_00] I think probably, yes. Like, one project which is kind of surprising and kind of mocked at one point was OpenAI's pen. You know, they had this pen thing so you could, when you write it, take down your notes and so on. This is an interesting concept because in some way it's seen as being backwards. It's like, why would you create a pen? And the other side, a lot of people still write a lot of their notes on paper. I'm the same.
[00:19:48] [SPEAKER_00] Like, even to-do lists and stuff like this, I still, there's something to be said about writing something to formulate your ideas better than typing it in the computer. And sometimes I still do it via pen or paper and pencil. And I think that's interesting. Then you have, I think, a lot of these collections, for example, the Bodleian Library in Oxford. It's one of the biggest archives of research material.
[00:20:15] [SPEAKER_00] I have not confirmed it, but I have no doubt that OpenAI are looking at trying to get access to that because it is, like, literally a treasure trove of information and books and knowledge that many people haven't seen. Which in that case is kind of a shame as well, right? The fact that it is locked away is kind of a shame. You know, Google Books, when it came out, Google Books was amazing. On one side is Google, so maybe I have to say it's not great.
[00:20:45] [SPEAKER_00] But on the other side, what they did was incredible because, you know, there is this, you know, having all this knowledge, but in a book that you had to go to a library, and that was nice in itself. But to find the information was really laborious, and you had to go, it takes a long time to be able to find the right book and the right information, and then laborious to go to each shelf and take off each book. And now just having this tool where you could go and search across all this knowledge
[00:21:15] [SPEAKER_00] and find the references that you're looking for. I remember at the time when I was at Oxford, you know, in the colleges, you sit beside, like, everyone at the table, like, different profiles and so on. And there was one girl who was doing her PhD on English literature, but looking at how electricity, the advent of electricity,
[00:21:43] [SPEAKER_00] became evident in literature over time. And she was spending a lot of time in the library, and it was kind of Google Books was around, but it wasn't super well known. And I told her, have you looked at Google Books to see, you know, type in electricity and see when the trends are, but when it first exists, first is mentioned and so on. And she said, no, I never heard of it. And then I showed her it. I showed it to her. She said, wow, this is amazing.
[00:22:12] [SPEAKER_00] And okay, it doesn't have everything inside it, but it opens up, you know, the ability to type a word and then see how that word was trending over time. It would have taken her years maybe to go through the same collective knowledge to be able to have something similar.
[00:22:30] [SPEAKER_01] Yeah. And one last question. Are there machines that have been produced that can turn pages and scan? And there must be some robotic, yeah.
[00:22:45] [SPEAKER_00] Yeah, they have. I mean, I know that some of them are quite quick. So you've got ones which are basically like a magic wand. You pass over the page. So you pass over the page, you pass over the next page and so on. And then I think there's a way then that they can control, like, to understand the context about what page it was. Yeah. But if you read it properly, of course, you have the page number as well. Yeah. But, you know, this is something that's been done, you know,
[00:23:16] [SPEAKER_00] automated by Google already for Google Books. And then existing systems, like, I don't know exactly what they were doing at Anthropic when they did the same thing. What I understood was that they basically tore out the pages. But that seems a little bit brutalist for my liking. I would have thought that, you know, this tech company would have come up with something more interesting. But maybe Occam's Razor came to effect here and the simple solution wins.
[00:23:45] [SPEAKER_01] That's right. Literally a razor, right?
[00:23:48] [SPEAKER_00] Yeah.
[00:23:49] [SPEAKER_01] Yeah. So, yeah. So let's talk about, before we get again to Born Private, the model that you're using, that you've built, where the proton is built, if there's no persistent memory, doesn't that make it very limited in its use?
[00:24:15] [SPEAKER_00] So just to clarify, we don't use our own model. We don't build our own models. The reason is primarily financial. It would cost us too much money to do that. Instead, we use open models, or different flavors of open, I would say. So, like, Almo was one that we were using at the start, Almo 2.
[00:24:38] [SPEAKER_00] We're not using it anymore because it's not, you know, since July last year, where we launched Lumo, the expectations about what these models can do is changed somewhat, I think, over that time period even. So we're basically keeping up with the frontier models by deploying the best current open models that we can deploy.
[00:25:04] [SPEAKER_00] So that includes things like GLM 5.1, KEMI 2.6, which is released just today. You know, QEN 3.5, we were using GPT-OS as 120 billion. But basically, any model that comes along which looks to be performant, like the next frontier in performance, we will use that. Nematron is another example. Nematron by NVIDIA.
[00:25:33] [SPEAKER_00] They have released two of the three flavors. So the first flavor was, I think it's, I can't remember exactly if it's 20 billion parameter model, a 30 billion parameter model. They have a 120 billion parameter model, which they just released recently, and they will have a 500 billion parameter model as well. They're using fully open code, open data. It aligns very much with what we want to do as well.
[00:26:04] [SPEAKER_00] If that model is close enough in terms of performance, we'll use that too. Then there's a Pertus, which is the Swiss AI one. The version 1 of that was probably not at the level where we could deploy to users yet. It wasn't created for that purpose, I should say, also. It wasn't created as a general-purpose chat GPT type of competitor.
[00:26:29] [SPEAKER_00] It was created as a base model that was going to be used for different applications. So in sciences and, you know, in legal context and financial context and so on, that could be fine-tuned towards those different application cases, but not as an out-of-the-box chat GPT competitor. But their upcoming releases of their plans will also improve that as well. So if it's close enough to something that we can deploy to production that users like,
[00:26:58] [SPEAKER_00] we will be able to run that as well. And for the context side of things, it's more complicated, let's say. So we don't, just because we're end-to-end encrypted for the chat history, for example, doesn't mean that we cannot use the information that's already there, right? So there is no real persistent context. There's different flavors of that. So, for example, memory. Memory itself is, it can be implemented in many ways.
[00:27:26] [SPEAKER_00] It can be like a global database that we are pushing to all the time. It can be something which is local to the user that you can just look up as you're making calls. It can be something which is hybrid. We keep it locally and then we sync it every so often, which is basically what we do. So when we sync, we keep it locally, it's encrypted. When we sync it, it's encrypted with the user's key, so we can't read anything on our servers.
[00:27:54] [SPEAKER_00] For Drive, for example, in projects, so Lumo has this projects feature, you can link your project with a Drive folder. And then that Drive folder can have all the information that you need to be able to answer questions in that domain. So if you're an enterprise and you've got one team which is focused on some financial stuff, you can have all your financial documents loaded there. What Lumo does then is it goes and looks, that folder is now synced.
[00:28:23] [SPEAKER_00] It downloads all those files. It transforms them into text. We index them all locally. Locally. So then we know for, we can basically do two types of search. One is just basic keyword type of search. But the second is that you can, you know, you can type in full prompts and it finds the most relevant documents based on the most relevant extracted terms. And then we inject those documents automatically into the context for the prompt
[00:28:50] [SPEAKER_00] and then send those back to the GPU, get the result back, and then the user has their context from their businesses within the response because we provided it at runtime. So we do, even if I think doing things properly in end-to-end encrypted environments is hard, but just because it's hard doesn't mean we can't do it. So we've already, we've been doing it for mail.
[00:29:20] [SPEAKER_00] I think search is particularly hard because you need the data on the client and not all clients are at the same level and the ability to be able to take all that information. With projects, we sort of, we made the decision that if someone links their drive folder, they would link something which is more around the topic that they're looking at rather than say link all my drive files, which is basically could be everything.
[00:29:48] [SPEAKER_00] It would be a lot to be able to sync onto your machine. Whereas linking a folder, which is very narrow or much more narrow, it gives us the ability that we can do that quite well without overloading the user's machine. So we do have ways. And then you've got, you know, the whole search thing. Our search is pretty good. So we have web search. We have financial searches with APIs in the background.
[00:30:14] [SPEAKER_00] Users can enable or disable those as they see fit depending on their threat model. So if they think that they, you know, I don't want the risk of anyone seeing anything I'm typing and any query going to some third party, you turn it off. Right. And if you do want it on, you can actually see what it's searched for and what results it got back and so on as well. Yeah.
[00:30:39] [SPEAKER_01] So Born Private. Describe born private to me, to listeners. Describe born private.
[00:30:44] [SPEAKER_00] It's basically, I mean, the basic premise is more that parents can go and reserve their email address for their kids. I would say, you know, the process is that you choose your email address, you donate $1 or whatever you want to the Proton Foundation to support the privacy mission. It's more of a symbolic donation. You get a secure voucher back to your, basically a link.
[00:31:13] [SPEAKER_00] And if you want to unlock your account for your child immediately, you can. If you want to wait for 15 years, you can as well. That's it. It's basically, I would say, born private more than being that particular feature, which, I mean, maybe it sounds underwhelming at this point, is more about the act of thinking about your child's privacy. Where does the privacy journey start? Or where does the data collection journey start?
[00:31:42] [SPEAKER_00] Depending on your point of view. You know, so some of the things that we like to talk about is that, you know, once a parent starts having emails with their gynecologist, for example, or perhaps for a fertility clinic or for something else, suddenly these systems outside are able to know, okay, these people and I want to be parents. We should already start sending them advertising on maybe some hospitals for,
[00:32:10] [SPEAKER_00] you know, that are good for birth or some advertisements for some gynecologist or whatever else, or pediatricians and so on. And as soon as people start thinking more about how their profiles are being created and how the profiles of your child is being created even before they're even born, I think that part is the most powerful part of the initiative, really.
[00:32:38] [SPEAKER_00] It's more about, you know, thinking about your child's privacy, thinking about the privacy of not just yourself, but also the people around you. And also highlighting some, I think, important things about what tech companies and what companies in general are doing, which is not at the benefit of you at all, but only at the benefit of them in advertising.
[00:33:01] [SPEAKER_01] Yeah. Well, maybe you can explain how that data is collected and merged, because we've all had the experience. You know, my wife is convinced that her iPhone listens to the conversations because we'll have a conversation at dinner and then she'll get served an ad. But my argument is, no, your profile is very complicated.
[00:33:29] [SPEAKER_01] You interact with digital systems all the time, and it's all companies, data brokers collect all of that and create a profile. But how, I mean, it surprises me that Google, who, of course, if you have Gmail, has access to your emails, uses that data, the contents of the emails.
[00:33:58] [SPEAKER_01] I know they use the metadata. But can you talk about that ecosystem and where all of that data gets collected and where all of it goes through email, I mean?
[00:34:14] [SPEAKER_00] Through email? Specifically through email? I mean, from...
[00:34:18] [SPEAKER_01] Or generally, I mean, that experience that we've all had of talking to your spouse and then suddenly getting ads served related to that.
[00:34:31] [SPEAKER_00] Well, Google is a complex machine, because not only do you have your Gmail, you have your Google Home system, perhaps, that's listening for keywords that are going to come up. Like Amazon Alexa has plenty of examples of them listening also to conversations of people as well. I think there's a lot of conspiracy theories about what's done and what's not done. And those conspiracy theories happen
[00:35:00] [SPEAKER_00] because there's a lot of gap in knowledge, which is also intentional, right? There's no real transparency over to say how exactly these profiles are created in the first place. But let's say from a very simplistic nature and only focusing on email, you create an email address, and the first thing you do is go off and sign up for Instagram, right?
[00:35:26] [SPEAKER_00] That tells me already something about your age, probably, let's say. Okay. Then you go in and sign up for some newsletter, some political newsletter, right? So now I know basically your age and probably also your political leanings. Then you sign up for some book or newsletter on AI, right?
[00:35:55] [SPEAKER_00] So at this point now I've got three data points, and then from that I can infer a whole bunch of stuff, right, just by connecting those things, like which people are interested in AI and also have right-wing ideology, for example. And then you find a whole bunch of stuff, which people are using Instagram or the age of Instagram users, let's say, also interested in AI and also are interested in this right-wing commentator or left-wing commentator or whatever else. And then you have a whole,
[00:36:25] [SPEAKER_00] you start off with those three data points and then you end up with a whole bunch of suggested data points. And then if you're a place like Google, you can basically say, oh, let's send an advert for this thing that's tangential, right? We think you might be interested in this. And then the user clicks on it, right? And they say, okay, the user was interested in that. Now you've got a new data point that helps you build a different profile. Or you might create a, send another advert and they don't click on that for things which are not super clear. For example,
[00:36:56] [SPEAKER_00] maybe you don't have any information about their political leanings or say religion, let's say. And you send an advert for some Catholic church thing, whatever. Maybe you're not going to click on it anyway. I don't know. But maybe you don't click on it, but then you send another one for a Protestant one. And you say, are they clicking the Protestant one? Okay, now we know their religion too. So Google and other providers like this, which it's not just about creating the profile with what data you give them directly,
[00:37:26] [SPEAKER_00] but also how they change. They're able to sort of interrogate certain questions that they might be having, their own model gaps, by basically proposing content and seeing what you click on. Or say putting a link in a Google search result higher up or lower down and seeing if you click on it or don't click on it. So changing all these visual cues and changing everything around the version of the internet, which is being tailored for you,
[00:37:52] [SPEAKER_00] also can have an impact on how you perceive the world, right? So let's say you didn't have politically right-leaning ideologies, but then there was an advertisement came up, which was super interesting and something you agreed with. And all of a sudden, you're pushed down a rabbit hole of believing in this ideology, but you didn't have that inclination to start off with. And that's even a more interesting thing because you're starting to change people's behavior
[00:38:23] [SPEAKER_00] and change, not only in changing who those people are based on what you serve them and so on. And I think there is an interesting film. I met the director, in fact, last month it was. The director is called Mark Silver and the film is called Molly Versus the Machines. I don't know if you heard this story about Molly Russell. She was 14 years old. She killed herself. Things, severe depression caused by,
[00:38:52] [SPEAKER_00] basically she had suicidal thoughts and the Instagram feed kept propagating more and more adverts and more and more content, let's say, on suicide. suicide. And in the end, she killed herself. And the whole thing is about, you know, you think your child is safe in their bedroom, but all these companies are basically changing the behavior of your of your child, even if you're not in the room with them and they're not physically in the room. But,
[00:39:21] [SPEAKER_00] you know, there's no real safeguards to stop people having a negative impact on your and your child's psychology.
[00:39:29] [SPEAKER_01] Yeah, that's that's fascinating. And the way you describe what hadn't occurred to me this this profile that different companies are creating of you is is kind of a cloud that's constantly morphing through time. Right. It was reminds me a little of your visualization work. So how does how does Born in Private work?
[00:39:56] [SPEAKER_00] Well, Born in Private starts by saying, you know, a lot of people's their digital identities are anchored in the email address. Right. So why you're getting information from at the start. So if the parents already are using Proton for all the communications with their their their gynecologist and hospitals and whatever else, basically none of that information is ever going to be used for any of to create any of those profiles. Now, that doesn't stop people from
[00:40:26] [SPEAKER_00] you know, if you use Proton but then, for example, use Google Chrome and you use Google Search. Well, you've basically negated a lot of the benefits, right, from using something like Proton because on one side of things, yes, you don't have all these things that are coming into your inbox that are tracking you, making sure that you're seeing what advertisements you click on. We're not profiling at all, right? So, you know,
[00:40:56] [SPEAKER_00] Google already knows that you're going to see your gynecologist therefore, or the, and maybe you've got some apartment for some fertility clinic, then it already knows that you're planning to do those things so it can start creating advertising profiles. If you start going and doing that also on Google or on Google Chrome and so on, then they will know that information too. So, you know, Born Private is more about, yeah, this is one step in the process but it's more to highlight that
[00:41:27] [SPEAKER_00] we should be trying to increase the sphere, the privacy sphere for individual users but also within that individual user a lot of people make decisions which are also going to impact their entire family tree, right? So, you know, a classic example is like if you go to 23andMe a few years ago and got a genetic test, not only have you basically made a privacy decision on your behalf
[00:41:56] [SPEAKER_00] to give away your data which can be used for something in the future but also for all your kids but also your family, your brothers and sisters and aunts and uncles, you've basically given away part of their privacy as well. understanding basically if Born Private does anything, it's about trying to increase the understanding of what privacy means, where it starts. Reserving an email address is one part of that but it gets the conversation going in people's heads,
[00:42:27] [SPEAKER_00] like how you want to deal with your kids' privacy online. So, I have a daughter, there are no pictures of her online. You know, every time the school asks, the school does ask at least, do you want the photographs to appear and are even internally as they know. My family tried to, my sister for example, will try and take photographs and put them on Instagram or whatever. I say no,
[00:42:56] [SPEAKER_00] you can remove her from the picture. We might appear in the pictures because I think our day is gone probably, but for her we've made the explicit decision that we don't want anyone that we don't know seeing pictures of our daughter. And if you see what happens, you know, like later on with deep fakes of the potential features for how all that information can be misused,
[00:43:26] [SPEAKER_00] I think people would take a second look at their decisions and maybe not do that.
[00:43:32] [SPEAKER_01] Yeah. Do you think that as awareness of privacy and the negative impacts of lack of privacy spread through society, that they will develop an ecosystem around, maybe Proton will be the kernel, around an ecosystem of privacy to counter the Googles and metas of the world so that,
[00:44:03] [SPEAKER_01] you know, someone with a born private email address could then participate in social media or participate in search without having them, you know, having somebody build this elaborate, detailed profile of them.
[00:44:24] [SPEAKER_00] Mm-hmm. So, I mean, at the minute, I would say maybe there are private social media systems, but I'm not sure if they're any good or if anyone's using them. You know, platforms like Facebook and Instagram and these types of tools, they're basically a TikTok or whatever, they're basically the dominant force and many people
[00:44:54] [SPEAKER_00] have made the decision that either they don't recognize the risk and their decision making has been basically zero. They just fell into the crowd of peer pressure. You know, your friend does it, therefore I have to do it too. If you don't do it, then you're left out. I think a lot of people are in that boat. I think maybe social media as it stands today will not be the social media that stands in a year from now or five years from now,
[00:45:23] [SPEAKER_00] especially with all the pressures that are coming on regulation for how people are treating people's data. And the example I give for Molly Russell, for example, is a big part of that type of drive, right? But people typically gloss over danger for convenience, but that's not something which is unique to privacy space. It's probably for almost everything, right? If something is easy, people are probably just going to do it.
[00:45:54] [SPEAKER_00] If something is a bit harder, then it's going to be about more resistance to doing it. It's easy to follow the crowd, it's harder to reject. I think with Proton, you know, before ProtonMail came along, it was kind of difficult to use BGP. I mean, for creating secure email, it wasn't that easy at all. And Proton came along and then made that whole process easy and transparent. The idea is to get to a point where privacy
[00:46:22] [SPEAKER_00] is not just the default, but it's also easy for people to do. And sometimes it's a long process because there will be unique cases, unfortunately, which show what can really go wrong. It's the same what you get from data breaches, right? It's easy to set up a website or to vibe code a project these days or whatever else. It's still hard
[00:46:51] [SPEAKER_00] to do security well. And so people just ignore security as the first thing. first stage, then they're breached and now suddenly all the records for all their customers or everyone that's been on their platform is now leaked online. And they say, oh, that was a mistake. I should have invested more time on the security side. People, there's drip thread, I would say, data leaks
[00:47:21] [SPEAKER_00] in the general population and people are more wary of it. For identity theft in the US, for example, is super easy if you have just a small piece of information. Less easy in Europe because you can't just create bank accounts so easily in Europe with someone else's name. But all these things are, if you make it easy for people to have a private
[00:47:50] [SPEAKER_00] ecosystem where they don't feel like they're left out because they don't use a certain tool, they're overall safer for it. They have no manipulation. You don't have cases where companies are trying to modify people's psychological profiles in real time. If we can get people to think more like this, that would be a success. Ideally, they'll come to Proton and use our products where possible. I don't see a day where we
[00:48:19] [SPEAKER_00] create a social media platform. Email is kind of like social media, let's say. You send messages, you've got connections and groups and so on. But I also don't think maybe social media is not going to be super popular in a decade from now.
[00:48:43] [SPEAKER_01] Yeah, because of all of the risks. But on the other hand, there is a positive side to this profiling because it does make you aware of things. You are served ads. of things that you are interested in that you wouldn't have found before. Is there a way of managing that so that it's not abused?
[00:49:12] [SPEAKER_00] Well, transparency is the important part of it. I mean, they do something on this, like you say, why did I see this ad? That type of thing does happen on particular platforms. But the regulation call, as it stands, is happening because we've long treated social media as some sort of utility function, rather than being a platform that
[00:49:40] [SPEAKER_00] was basically abusing the users, keeping them hooked, the same way as you have for cigarettes or alcohol or gambling. I think gambling is the best example, and I brought this up a while back. I think I have read more that people are using that example also in public discourse, but gambling systems have regulation in place because otherwise, what's to stop the gambling company from just taking everyone's
[00:50:10] [SPEAKER_00] money? money. And if you don't have these regulations that people abide by, then the population would be miserable. You'd have a bunch of people just taken for granted. They would probably be higher divorces and mental health problems and so on. But if you look at the effects of social media, it's largely the same. People who are addicted to it, they scroll all day because of the things that have been engineered to make people stay on the platform.
[00:50:40] [SPEAKER_00] The whole thing is about increasing active users on the platform. So increase engagement. The more engagement you have, the more time they spend on your platform, the more opportunity there is for shareholders and they will invest more so stocks will go up. Basically the reward system is broken. There is no real incentive for social media companies to try and limit usage because it would damage themselves. So I think when these types of things happen where the user
[00:51:11] [SPEAKER_00] health and the user focus is diverging from the platform goals, then you need regulation to bring them back together. And I think it's fair that people are talking about it now because for a long time people just said, yeah, let themselves regulate and so on. But clearly that hasn't worked.
[00:51:33] [SPEAKER_01] Yeah. And also our digital lives have overtaken so much of our real lives that in the beginning you didn't spend that much time in the digital world. But now we're unfortunately we're all doing what we're doing right now, speaking through a digital medium. The
[00:52:03] [SPEAKER_01] Proton, so Proton's kind of a Proton, born private is kind of a starting point. At least you know with a Proton email address that the content of your emails are not contributing to this profiling. Exactly. Yeah. How does somebody get and born private is a great concept but presumably
[00:52:32] [SPEAKER_01] anybody could get a Proton email address.
[00:52:38] [SPEAKER_00] Yeah, just go to Protonmail.com. Is it Protonmail.com? I'm going to check just to make sure. Yeah, Protonmail.com. It's been a long time since I actually had to go to the website. Yeah, Protonmail.com and create a free account or pay also. Paying is good. It gives us jobs and so on. But yeah, Protonmail.com for create your free account and that gives you then the access into the rest of the ecosystem. You've got
[00:53:08] [SPEAKER_00] access into VPN, to Proton Drive for your document storage or also photo storage. We have Proton Pass for your password manager. We have docs and sheets. We have Lumo for the AI assistant. Meet, which we launched, I think, two weeks ago. Two weeks ago, three weeks ago, which is our video conferencing solution. So you can keep everything
[00:53:37] [SPEAKER_00] private there too. And more. So it's pretty easy to set up.
[00:53:44] [SPEAKER_01] Yeah, that's fascinating. you can build, I mean, as you were saying, you're not going to create a social media platform, but you can build a digital ecosystem through Proton. At least within that ecosystem, your privacy is protected. Yeah,
[00:54:01] [SPEAKER_00] and we just released Proton workspace, which is also our plan, more for B2B focus. But the idea is that basically companies can go along and they will have everything that they need to be able to function as a company by using Proton products. So all of your IP stays with you. You don't have to worry about confidential data getting lost somewhere. If you want to use AI
[00:54:31] [SPEAKER_00] assistance, you can use it within ProtonMail. We also have Scribe, which is the first AI integration we did, which is similar to what you have in Gemini now for Gmail. And you can use Lumo for everything else and connect it to your documents. You can ask questions about it and use it for your productivity gains, hopefully. And the whole idea is to create this ecosystem or workspace where you don't have
[00:55:01] [SPEAKER_00] to compromise on functionalities anymore to be able to use something which is fully private.
[00:55:07] [SPEAKER_01] And for born private, that's also what parents can find on the Proton website. Is that right?
[00:55:16] [SPEAKER_00] Yeah, if you search for born private in your favorite privacy friendly browser, then you'll be able to find it. But basically Proton.me slash mail slash born private. Okay, well, why don't we leave it there?
[00:55:34] [SPEAKER_01] That's all fascinating. reading.

