#340 Steffen Cruz: Training AI Without Data Centres
April 29, 2026
340
46:25

#340 Steffen Cruz: Training AI Without Data Centres

What if you could train a frontier AI model without building a single data centre?

In this episode of Eye on AI, Craig Smith sits down with Steffen Cruz, co-founder and CTO of Macrocosmos, to explore a radical alternative to the way AI models are built today. Instead of billion-dollar GPU warehouses, Steffen is training large language models using idle compute from devices distributed around the world, coordinated through the Bittensor blockchain.

Steffen breaks down why the centralised data centre model is heading toward a wall. Projects like Stargate and Colossus cost tens of billions of dollars, and as appetite for larger models grows, the economics simply stop making sense. He explains how distributed training flips this on its head, tapping into surplus energy, underutilised GPUs, and even consumer devices like Mac Minis to train models at a fraction of the cost.

We also get into IOTA, Macrocosmos's flagship technology, an orchestration layer that takes compute nodes scattered across the globe and makes them act like a single supercomputer. No single device runs the full model. Instead, each one carries a small slice, a technique called model parallelism, and together they can train frontier-scale models that would otherwise be out of reach for startups, researchers, and enterprises.

Finally, Steffen shares what he's building toward: 70 billion parameter models trained at 10 to 20 percent of centralised costs, a two-sided marketplace for compute, and a future where anyone with a spare GPU or Mac Mini can earn passive income while contributing to the democratisation of AI.

Subscribe for more conversations with the people building the future of AI and emerging technology.

Stay Updated:

Craig Smith on X: https://x.com/craigss

Eye on A.I. on X: https://x.com/EyeOn_AI

Timestamp:

(00:00) Introduction: The Problem With Blockchain AI Projects

(06:39) Meet Steffen Cruz: From Subatomic Physics to Decentralised AI

(09:16) What Is a Bittensor? The Blockchain Built for AI

(11:53) How the Blockchain Actually Works: Registry, Clock, and Rewards

(15:08) Why Data Centres Are Hitting a Wall

(22:01) Distributed Training vs Federated Learning: What's the Difference?

(27:47) Train at Home: Turning Your Mac Mini Into a Passive Income Machine

(32:49) IOTA Explained: Building a Global Supercomputer From Spare Parts

(39:43) How the Network Scales: From 256 Nodes to Limitless Compute

(44:39) The Road Ahead: 70B Parameter Models and the Future of Affordable A

[00:00:00] So my name is Steffen Cruz. I am the co-founder and the CTO of Macrocosmos. I hold a PhD in subatomic physics in the University of British Columbia. So I was a physics researcher for the beginning of my career, and then I pivoted to A.I. when it became apparent to me that there's a lot of opportunity for scientists that are just about to graduate from their postgraduate studies.

[00:00:24] And there's an entirely new science sort of being born and developed in real time before us. And I think my decision is one that a lot of other people have made as well, where they want to go from being in what can feel like the sort of the old school machinery of traditional science in academia, where things move slow and your contributions can be very narrow and sparse, to becoming a pioneer of this entirely new exciting thing,

[00:00:52] which is not only abstract and theoretical, but it's being applied in new and interesting ways constantly. So when I saw that, I concluded my PhD and I decided to leap, leap, leapt out of academia and into sort of the application of A.I. and also physics in various private sector domains, such as manufacturing and systems optimization. Following that, I discovered BitTensor, which is a very interesting A.I. blockchain project.

[00:01:20] And ostensibly the purpose of BitTensor is to allow A.I. to be developed in a way that is fully democratized and is sort of done in a global way with incentives. And I thought that was such a fascinating proposal that I decided to join the network. I've been within the BitTensor ecosystem for around three years now, and I have been relentlessly experimenting and building within BitTensor

[00:01:45] the kind of things that I think would make a lot of sense and would bring a lot of value to the world using the blockchain as a prerequisite tool to enable A.I. research to be done in various ways and scales that will be very difficult otherwise. Yeah. And specifically, well, first, let's talk about BitTensor.

[00:02:05] Before you began, we were talking about how there are a few of these blockchain ecosystems for either training or marketing A.I. models or A.I. applications.

[00:02:23] And I'm familiar with SingularityNet. BitTensor is much larger. Is that right? Is it the largest of these or one of the largest and how many are there? I believe so. So BitTensor is actually over 100 projects under a trench coat. And I think that's also why it feels so encompassing and so large.

[00:02:46] And this is what is very interesting about BitTensor is it doesn't seek to solve a specific narrow problem in the A.I. industry. It's not like trying to solve something like how do we provide people with access to models or how do we train models or how do we do one of a plethora of different things?

[00:03:02] It actually has done a lot of foundational work on how do we make it possible to build a base layer for people to come and solve just about anything they can imagine in a way that uses the blockchain as a reward mechanism, as a coordination layer. But then BitTensor does a very good job of stepping out of the way and letting people's creativity drive them through to build what they want on top of it.

[00:03:25] So BitTensor feels very big and very sort of like it's everywhere because there are today there's 128 different teams that build on top of BitTensor and they're all building something different. A lot of them choose to work within A.I., but there's also some that work outside of A.I. as well. And so for that reason, I think that BitTensor punches well above its weight.

[00:03:44] Yeah. And for those who are not familiar with the blockchain and even me, the blockchain in effect is a registry of addresses, right? And it points toward the actual assets that are being registered.

[00:04:04] I mean, this is the thing about blockchain art that was such a craze for a while that you're not actually buying a piece of it. I remember Trump was selling these, I think, playing cards or something like that on the blockchain. And you're not actually buying anything. You're registering your ownership of something. Precisely.

[00:04:34] But the actual asset is a digital asset that can be copied millions of times. It's up to you then if you want to enforce your ownership to track people down and that sort of thing.

[00:04:49] So in the projects that I've seen related to A.I., it's a registry that points to an A.I. project or is there something more happening on the blockchain? So in a project like BitTensor, the blockchain is used in a number of different ways. But fundamentally, it's an immutable database that is globally shared.

[00:05:18] And that in itself enables a lot of things that Bitcoin pioneered some 15 years ago. What becomes possible when you have this sort of shared globally distributed state is that you can create a store of value and then you can start building higher and higher order commodities on top of that. So Bitcoin sort of addressed the most principle version of this. But what was demonstrated as a result of that is you can actually aggregate a huge amount of compute or let's just call it effort.

[00:05:47] Once you have this interesting distributed sort of system architecture, people come to it. And because there's economic incentives involved, people are constantly innovating and optimizing and they're trying to find a way to sort of percolate to the top and to become the most competitive people that are doing whatever this mining task involves. In the case of Bitcoin, the mining task is just guessing random numbers faster than anyone else can.

[00:06:12] That's why you just throw a more powerful computer at it and you become more competitive miner on the Bitcoin network. And other crypto projects have tried to orient that effort, that sort of collective coordinated effort of all the participants have tried to orient that towards different tasks, whether it's distributed file storage or creating a global residential proxy network or solving data scraping from the web. And all of these things are really interesting.

[00:06:38] Whereas what BitTensor uses the blockchain for fundamentally, as you described, the registry component is very important. It's also sort of a shared synchronization clock. The notion of blocks and the blockchain gives everyone sort of a drumbeat that you can actually anchor work to, which is an integral part of a lot of our research on BitTensor. We have this sort of shared clock, which the blockchain regulates.

[00:07:02] But ultimately, you can do as much or as little on-chain, as they call it, or off-chain as you want. And I think some of the benefits of doing things on-chain is you have complete auditability. You can see precisely what the details were that went into a specific piece of work or transaction, so it gives you transparency. The problem with that, of course, is if you've got billions and billions of transactions being written to a database, your database gets very big very fast.

[00:07:28] And then it becomes kind of unneulably for people to actually have a copy of, which can interfere with the true decentralization ethos of it, and you end up with just a few copies. So I think that there's a pragmatic middle ground where you store exactly what you need but nothing more to ensure the core parts of the work that you care about are captured and transparent for everyone to audit, review, and verify. Whereas the rest of it, you can store elsewhere. You can store offline. You can do a million other things with it.

[00:07:57] So I hope that adds a little bit of light to the question, Craig. Craig. Yeah. In your case with Macrocosmos, you're providing compute for training models. Is that right? Explain precisely what Macrocosmos does. Absolutely. So today we operate three subnets in BitTown Sun. Each subnet is, you can think of it as effectively a different project, a different service.

[00:08:23] And our three different subnets, I would consider them to be three different axes or three different use cases for the blockchain that I think are very potent and powerful. So one of them is just the aggregation of raw compute at scale. And if again, harking back to what I described about the Bitcoin network, I think one of the most important things of our time is set to be AI itself. And those who train the models and own the models will have an enormous advantage.

[00:08:52] And so a big part of the purpose of BitTensor is to provide an alternative way for us to create these models, to train these models, to deploy these models such that they're not held in walled gardens and we're having to pay for the privilege of access to them. But we want to reinvent that in a way that's more democratic and also borderless, something that's censorship resistant, something that cannot be controlled.

[00:09:18] So our primary, our flagship project within BitTensor involves what's called distributed training of large language models. And that's quite a mouthful. Distributed training is quite different to how a lot of the frontier AI labs train models today. Typically, you have huge warehouses stacked from floor to ceiling with GPUs, which are specialized computers for training AI models and serving AI models. And you've got hundreds of thousands of them, hundreds of thousands of them in the largest data centers in the world.

[00:09:48] And they're all plugged up together with extremely high data transfer speeds. And as a result, you can basically use the whole thing as a single computer. And this is fantastic and it's been demonstrated that you can just scale the amount of compute and the amount of data you throw at a model and they get predictably better and better and better. And this is called a scaling law. Now, there's an alternative to this, which is called distributed training.

[00:10:10] Distributed training doesn't rely on having a single warehouse stuffed with computers, but actually it's based on the premise that you can train an equivalent model using compute computers that are distributed around the world. And now this has a lot of very interesting outcomes. When you don't do everything in a single data center, one of the first things is that the capex of actually building this massive warehouse in the first place is massively alleviated.

[00:10:37] There's also the local energy that is required to run one of these massive data centers can also now be distributed around the world in a way that creates less impact on local communities, generally much, much less detrimental to the environment. And lastly, it allows you to do something that is actually fundamentally not available once you've built your data center and stuffed it with GPUs. So, the cost of training models is already baked in to that initial build process.

[00:11:06] Whereas in the case of distributed training, which is what we care about, we can actually perform what is effectively cost arbitrage. So if there's a surplus of energy in Iceland, very, very cheap energy, we can actually use that pocket of energy, even if it's only for 12 hours of the day, and we can target a lot of that compute and use it in a very elastic way. And in fact, that's what we do. Today, we are training models, we're training multiple models all at once using pockets of cheap energy, which translates into cheap compute.

[00:11:35] So it has all of these different economies of scale, which are certainly unusual, but are becoming more and more accepted as a viable alternative. And as our appetite for bigger and bigger models is only going to increase as the years go on, we're going to find that we need to start looking at alternatives before we hit a hard ceiling on a lot of this. Because things like the Stargate project and the Colossus project, these are multi-billion dollar GPU build-outs.

[00:12:01] And so we think that just like the fundamental physics experiments, like the Large Hadron Collider, at some point, you require the budget of a nation state, it's $20 billion to build a bigger ring to smash protons or electrons. You need to start thinking about different experiments, because it just becomes unpalatable eventually. So we're trying to get ahead of this problem. I think that by 2028, we're actually the world, the mainstream, the Overton window about training models is going to shift.

[00:12:29] And we are going to have to start thinking about how do we do this in a way that arbitrage is more efficient in both cost and energy. So we're trying to do a lot of the early work right now on that. And BitTensor is a wonderful place for us to do this research because we are able to use the blockchain to reward people that contribute their compute to our training experiment. Let me ask a couple of questions. I had a guy on the program who was doing something interesting.

[00:12:55] He was, and I don't remember whether it was blockchain based, but he would aggregate spare compute from data centers or, you know, whether on premise or, you know, independent data centers around the world,

[00:13:16] and then offer that compute to AI projects at a cheaper price than the hyperscalers or the big clouds could offer. And so is it, that's one question, is it like that and you're just tying it to the blockchain to organize it and distribute it?

[00:13:45] And the other is, you know, for a long time, people were talking about federated computing so that you wouldn't have to lose control of your data, but you could make it available for training models.

[00:14:05] And again, that was using the blockchain to register the data, I guess. So first of all, is it related to either of those concepts? And then the deeper question and the one even talking to the people I've spoken to, I've never really gotten a hold of. So the chain again is just a registry.

[00:14:35] There's no training data on the chain. There's no compute on the chain. The compute resides, as you said, maybe in Iceland. What is the link between the chain and the compute in Iceland? So starting with your first questions, the relationship between those two examples you gave and what we're doing, I would say that they're cousins, but they're not siblings. They're a little bit different.

[00:15:04] For example, federated learning is built on the idea of anonymizing client data, but being able to continuously train models. We were working with a slightly different version of this, which is not so much the edge devices providing training data today, but the edge devices are providing the trained compute. I suppose would be a way to think about that.

[00:15:28] And secondly, I think regarding what BitTensor actually does or what the role of the chain is, at its core, it's really not that revelationary at all. At its core, what the blockchain is effectively doing is it's providing a trustworthy, transparent record for everyone to audit, especially when there's code that's being run on the blockchain as well.

[00:15:56] People understand how this code is going to interact. It can't be tampered with, which means that it's predictable. And that's a proxy for safety and trustworthiness. So if someone has already deployed a smart contract, for instance, we were not actually building a smart contract in microcosmos, but I guess just to illustrate the point, if someone has written something in the smart contract, then the creator of the smart contract can't later change the terms and conditions on you.

[00:16:21] The point is, it's a modular block of code that will run under these conditional expressions. And that makes it that makes it something that you can effectively count on. So I believe I've given you a bit of an illustration of one of the projects that we're working on in BitTensor, which is distributed training. This project is called IOTA.

[00:16:40] And as I mentioned, I think that it's a precursor to some very powerful downstream technologies that we may find are an instrumental part of the way that we create AI in the next decade, let's say. Yeah. We really do believe that there's also a real change in the way that people are thinking about personal agents and personal compute. So I've seen personally in the last few months, a lot of people are starting to stockpile Mac minis.

[00:17:10] I'm not sure if you've seen this trend, but basically now that agents are starting to become kind of economically useful a little bit, we're seeing the beginnings of it, I think as a fair appraisal. So what we're starting to see now is people want one of these things running for them privately around the clock day in and day out. So they're buying a dedicated computer. And on that computer, your personal agent is running and it's checking your emails. It's maybe doing some online shopping. Maybe it's doing some work for you on the side or some hobbies. And this is a wonderful use case.

[00:17:38] And this is also a really interesting place where our technology of IOTA developed by Microcosmos comes in because all of these computers that people have now built at home. Well, there's two ways to think about it. That's your personal agent that is creating value for you personally. But wouldn't it be great if that compute could earn passive income? Almost like you own a property and you can Airbnb it out when you don't need it.

[00:18:00] And this is actually what we're offering as we develop IOTA is we've created something called Train at Home, which is people that have unused devices at home. Let's say MacBooks, Mac minis, or even consumer GPUs. They can plug those into our network. And that means that they become part of this global supercomputer. So we can use them for training models, which we can commercialize. And they can also basically create some passive income just from having that device sitting around that was otherwise underutilized.

[00:18:30] So I think there's something very parsimonious about that arrangement of things. And the more people that are choosing to buy personal devices to run these agents, you realistically don't need the agent running 24 hours a day. You need pockets of productivity. I want to make sure it, you know, plans my weekly agenda. I want to make sure it checks this, that and the other. Maybe there's only four hours of the day that you actually need this thing.

[00:18:50] And you can actually get a return on your initial investment within, you know, 12 days, 30 days by plugging it into a network like this and also being part of something that I believe has a great purpose, which is the democratization of AI itself. So it's a really interesting intersection of timelines right now with agents becoming demonstrably more useful and engagement increasing. And we basically can use all of that consumer compute. We can glue it together and we can train models at greater and greater scales and greater and greater utility.

[00:19:21] So, yeah. Yeah. Yeah. Yeah. As a matter of fact, I'm running the Mac mini stockpiling, I think started with the open claw and I'm running open claw on this computer, which I know is a bad idea. And I have an unused Mac mini upstairs. So I'll probably go fix that after this call. But the so.

[00:19:49] So how does somebody use this? Well, first of all, yeah. Back to the point of. So you've got compute around the world that is plugged into the network. Do people who are using Macro cosmos, do they need to know what's how the what the link to the blockchain is?

[00:20:19] Absolutely not. Yeah. And and the blockchain is just keeping track of where the compute is and what's available at any point in time. Is that right? Yep. And it's just responsible for rewarding participation. So it's making sure that any contributions made by anyone's devices are fairly, fairly rewarded. That's really the way I think about this.

[00:20:47] The blockchain is just a really convenient way of taking care of compensating people for their contributions. The actual onboarding experience for people is we keep this as low cognitive load as possible. Right now, it's a one click app store download. This thing will just run passive in your machine, adding some more progressively more useful user experiences, things like telling it, oh, you can only use my computer when I go to sleep. So it's by default, it's disabled until 10pm.

[00:21:16] And it's only enabled from 10pm to 6am. So that will be a simple user control that will, you know, get it out of the way of when you're actually trying to be productive with your machine. That's assuming it's your primary machine. If it's a secondary machine that just sits at home. We're also creating a way for agents to decide, oh, I finished my work, I'm not going to need to do anything for four hours. So the agent just decides, hey, I'm just going to go and make 20 bucks. And that would be nice to come home from work too, right? It'd be nice to come home from work and you're like, hey, what have you done today, Claude?

[00:21:44] And it's like, oh, hey, I finished all my work by 9.15am and you were going to be out till 5pm. So I thought, well, I thought I'd just shop around and see how I could make you a little bit of sidecash. And here you go. So I do think that as our computers become more autonomous in thinking about our relationship with computation, it's going to change fundamentally. And our computers are not going to sit there waiting. They're going to be very proactive and it is very resourceful. Especially when you've got these clever, clever little agents that are running the show inside of the machine.

[00:22:13] And as a result of that, I think we need to really rethink what your computer can do for you. You know, our relationship to that world. And I think bringing this back to the top about BitTensor, you can think of BitTensor as just a bunch of places where you can use that machine and you can go and make some income from it. Yeah. Whether it's providing something like human intelligence or whether it's training models or whether it's detecting if an image is real or fake.

[00:22:42] You can almost think of it as a mechanical Turk for agents and humans alike. Well, you can go there, you can monetize your skills and your experience or just your raw compute. Yeah. And for the training, how does, let's say I have a large model that I want to fine tune. How do I do it with Micro Cosmos?

[00:23:09] So today our work is focused on the first stage of training models specifically, which is called pre-training. So pre-training has historically been by and large the most computationally expensive part of creating a model. It's the part where they systematically inhale the entire internet, which usually takes months and tens of thousands of GPUs. Beyond that fine tuning is comparatively a much shorter computational task.

[00:23:35] So I think the, it's the essence of pre-training requiring long time horizon workloads. That is actually why we're so interested in moving this from a centralized to a decentralized context. Um, and there's a lot of pioneering research that needs to be done that we hope will be very useful even outside of the web three world.

[00:24:00] We hope that this is really useful research to anyone out there that is interested in economically or capital efficient ways of training models. We hope that this is a contribution to the field. That's really how I see this because what we're doing at the root of our work is we're taking a network of very unreliable compute. People are joining people are leaving. It's just constantly churning over and turning over.

[00:24:24] How do we create something that is persistent and stable out of something that is fundamentally so noisy and unstable and this, and it's, and it's core. That's the problem we're trying to solve. And it just so happens that the, the, the, the, the use case for this persistent compute fabric we're creating is trading models, but there's nothing stopping that from doing something else. Perhaps there's a world where universities could use a similar decentralized network for their traditional HPC cluster jobs where they're doing academic research.

[00:24:54] That's not necessarily AI related at all. It could be bioinformatics. It could be physics. The same thing, the same thing is still, uh, still remains to be true, which is if you need 10,000 nodes of compute for 12 hours, what's the cheapest way you can get them? And we're trying to build software that basically orchestrates compute around the world to act like it was all plugged into itself and it's a supercomputer.

[00:25:18] And I think it's a, it's so much, it's, it's a fundamental thing that we're working on that we're applying towards training models, but I hope has a much broader value to the community. Right. And so on pre-training, uh, how does somebody use it? Absolutely. So we're sort of emerging from, from research mode right now.

[00:25:36] We spent about nine months on IOTA so far, and we've demonstrated that we're able to reproduce a lot of important baseline benchmark performance metrics using this, what we call it the wonky vegetables, sort of the ugly vegetables that you get in a vegetable aisle that no one necessarily wants to buy. We're going to turn that into a premium soup. And we've basically graduated from research mode right now, where we're showing that the two are indistinguishable if you treat them in the right way, if you use the right spices, tastes just as good.

[00:26:05] Um, what, what we're moving to next is thinking very deeply about what is the, what is the commercialization opportunity for a technology like this? Who are our target audiences? We think that the world is going to train more models in the next 10 years, not less. We think that researchers, small startups, especially cash strap startups, academia, all of these people are interested in training more models, but perhaps for budget reasons, it's very hard for them to do so.

[00:26:31] So we're sort of targeting the commercialization of IOTA as two things primarily. One is what we call supply side. So supply side would basically mean people that have surplus GPUs. So the neoclouds, the hyperscalers, perhaps they have 10,000, uh, GPUs in their inventory, but they can only rent out 9,000 of them at a given time or 9,500 of them.

[00:26:55] Well, for us, them plugging in their compute into our network allows them to increase their overall utilization, which goes straight to that bottom line. So they'll be happy with that. And even better that if you actually to look at the usage of all of those GPUs that a hyperscaler has, you might find that, well, it's, it's rented out for four hours and then it's not rented for two hours.

[00:27:17] And then it's rented out for four hours. And that little two hour gap is exactly what we're trying to utilize these short interruptible bursts of compute. And we turn that into, again, a continuous stream of compute is exactly what excites us. So for the supply side, it's basically anyone that has GPUs, what they have to do right now is if they can't rent it out, this sort of a buyer of last resort on the market, they'll, they'll rent out those GPUs at cents on the dollar for inference tokens.

[00:27:45] And we can basically say, well, Hey, instead of doing that, how about you plug them into our network and we can basically give you a better margin. Because trading is a higher order commodity than inferencing. And as a result of that, we believe we could pass on the benefits to their suppliers. So that's the supply side on the demand side. As I mentioned before, we have researchers, we have startups, we have all kinds of different profiles of users that we know right now, their training models all the time.

[00:28:11] What we want to do to that for those guys is basically provide them with an interface to train models in a way that's recognizable to them. So there's already very popular libraries like PyTorch, like TensorFlow. These are very popular libraries that are basically used by everyone in industry that's training models. Well, we want to make it as simple for them in terms of abstractions as if they were just building models in the normal routine way.

[00:28:36] So with no additional cognitive overhead, you basically say, I want to train a model. You're not going to painfully specify, oh, I want to use a Nigerian GPU for 12 minutes. Of course you're not, right? What you're going to do is you're going to say, this is my, this is my sort of executive level objective. This is what I want to get done. These are the parameters that I want to have precise control over for reproducibility purposes, for deterministic purposes. And all of this section down here, feel free to go and arbitrage and get me the cheapest version.

[00:29:05] And then we have this real, now we have a two side of market, right? Now we have a sort of supply side and a demand side that are sort of dancing together where we're arbitraging, offering cheaper rates. And IOTA is a technology that basically is this infrastructure layer that is, is the one that brings this compute to market, pairs it up with users that want it. And we believe that that's something that is phenomenally valuable. Yeah.

[00:29:29] Well, and then how does the, the under the hood, how do you get the, the, the compute to the, to the model? Cause they're not sending the model around all over. Uh, yeah. The way that you get the compute to the model is we it's, it's, it's effectively a, an orchestration deployment layer.

[00:29:55] Um, something reminiscent of Kubernetes, but Kubernetes that deploys to heterogeneous compute nodes around the world. So it's, it's a dynamic deployment layer that basically takes some containerized code that the user writes, sends it out to those machines. And then very importantly ensures that those machines around the world have a tunnel to communicate with each other. And that tunnel is the part that makes them act like one big blob of compute instead of 50 different random pieces of compute.

[00:30:21] So it's, it's effectively that it's an, it, it deploys anyone's code as an interactive system or a network of compute that can, that is sort of real spatially aware. Right. And, um, and that orchestration layer, um, is, where is that based? The orchestration layer is something that we maintain and operate. Right.

[00:30:47] So you would effectively dispatch your code from your laptop or anything like that. And your code would then be sent securely to these backend nodes, wherever they might be distributed around the world. And then you would effectively, as a user, you would have the ability to monitor logs, monitor trace data, the usual sort of dashboarding tooling, things like that as the user. But effectively what's happened is your code has now been sent out, cloned replicators for all these different copies.

[00:31:14] And all of those copies are hosting a section of the model, which is what makes IOTA very, um, in my opinion, very novel. We don't actually have each of these compute nodes running a full copy of the model. They're actually running a small sliver of the model. It's called model parallelism. And in effect, what this means is you can train really large frontier size models using very small building blocks. Like you can build a huge Lego tower out of small pieces. It's the same idea.

[00:31:44] And the way it works is in order to train the model in its entirety, you have to root information from all of the sections of the model and knit that together. And as a result of that, you can basically train this, this sort of distributed model architecture as if it was all stuck together in one place. And that's, uh, that's why it's taken us nine months to get out of research mode, to be honest, Craig, it's, it's quite a formidable task. Yeah. Yeah. And, and is the blockchain important for the synchronization?

[00:32:13] Because you need a very tight synchronization between, uh, nodes of, or between, um, uh, different compute. Right. So the way that we utilize the blockchain today for IOTA's purposes is we need a registry of everyone that's in the network right now. So it exposes them as, Hey, I'm an available worker. Here's how you can find me.

[00:32:40] And here's my unique, uh, address that I know to identify you unambiguously. So we, we use that as sort of, uh, uh, an authorization layer, an identity there. Also with that, we, we have an off chain layer to IOTA, which is constantly tracking the contributions of that node. And then once we have computed the total work done by each node that is sent back down to the blockchain layer, where it's, it triggers a payout loop.

[00:33:10] So everyone that contributed that compute is then rewarded in the IOTA token for their contributions of work. And, and this actually can be articulated in a very tokenized context or even a U S dollars context. So all you effectively need as building blocks here is, is, uh, you need to know exactly who it is that's assigned to this model training experiment. You know where you can find them, you know how to glue them all together and you know how much work they've done.

[00:33:35] And if you've got all of those things that you have an operational distributed trading network, how large is the network and how do you, how do you grow the network of compute? Yeah. Um, so when we began the source of all of the compute for our distributed trading experiments was people that came to our project in, in that 10 sub the IOTA project with, uh, they have their compute.

[00:34:04] They know what we're here to do. Usually there's a philosophical alignment, or sometimes they just see it as an opportunity to make some cash, but basically people show up armed with usually pretty powerful hardware. So the actual, the, the, the, the supply side of computing bit tensor is very much not a concern. People know that bit tensers there people know the rule, you know, the rules of the game. So we, we, we, we almost always oversubscribed in terms of people that wanted to contribute to our experiments.

[00:34:31] However, in December, November, December last year, we wanted to go beyond the scale of the network that is supported natively by bit tensor. So the maximum number of people that could be participating in the IOTA experiments was 256, which if you consider that to be a data center, it's tiny. It's, it's, it's a pebble.

[00:34:52] So we actually worked around that and we've now made it, uh, effectively a limitless size system by instead of having every unique entity as a unique address on the blockchain, we now keep track of it off chain effectively. And what this allows us to do now is we can have an arbitrary size registry of participants. So we had 2,500 people download our Mac OS app in the first two weeks since it launched.

[00:35:18] And now all of those people are available at a moment's notice to join our network and they're paid out identically to the ones that are strictly on chain within bit tensor. We have a transparent payout system that people understand it's a predictable system. It rewards, it directly rewards contribution in proportion to basically hours of work done per day. But the nice thing about this is it's not limited by the amount of storage capacity for the blockchain or the number of slots that a blockchain can have as miners.

[00:35:46] We actually decided to just rewrite a lot of that using different rails so that it's not, um, it's not limited in that way. And now as a result, we have 500 or so, uh, today running that are the amount of work that we have to give the miners to do is actually quite, um, um, it, it changes over time. There's times when we run lots and lots of experiments. We train lots of bottles and there's also time when we only need a small amount of compute and we only pay for what we need.

[00:36:15] We're not just the force that isn't just always running. We actually have experiments that are live or we have production runs where we're going to say, this is the model we're going to train. We're going to go find 500 devices around the world. Fortunately, again, the, the, the, the bits tensor pieces made it very easy to find 500 random people in the world with powerful computers that I can just plug into my network, which is the miracle of technology really. Um, and, and so we can effectively not worry at all about where that compute comes from.

[00:36:42] And then downstream, we basically pay everybody out for that work that they've done. Yeah. And is there, do you have a sense of scale, uh, comparative scale that, that would give me and listeners an idea of, of how this would compare to, uh, to a major data center? Or, or. Yeah. Yeah.

[00:37:09] The, the largest data centers in the world are measured that the ones that have started to come online are now hundreds of thousands of GPUs. So these are multi-billion dollar projects. They're absolutely enormous. We're not, we're not at that scale yet, but the nice thing is the technology that we're building scales in a very beautiful way. Um, in fact that, you know, to go from a day to go from our network with 10,000 nodes to 20,000 nodes, there's nothing that we fundamentally need to change.

[00:37:36] We just need to expand the distribution surface of what we're doing and our infrastructure layer scales, uh, to much, much higher capacity. So again, this is the nice thing about not being committed to a bricks and mortar, um, build out in the way that the centralized compute paradigm is in our case, if we want to double triple or 10 X, the capacity of our compute network, we simply need to go out, increase the surface of the network. Yeah.

[00:38:01] So, uh, what I think is in the books for us is by the end of this year, I want, well, by the middle of this year, I would like us to reach 5,000 compute nodes. I think this is a respectable size cluster that you can train a model that will get people's attention and make them understand that there's a lot of utility in this technology. I think that's an important milestone for us.

[00:38:24] Something around the 70 billion parameter model size, I think is when you go from, um, go from, you know, you graduate from the sort of small models that are proof of concepts to things that are actually taken much more seriously. And enterprise customers that have been waiting to train an in-house model, whether it's a legal specialist model or whether it's a medical model or something like that, perhaps they've been waiting for a while, but they couldn't really afford it because it's a very high price tag that comes with training a model of that scale.

[00:38:52] Because of the discounts that we can offer using this energy and cost arbitrage approach. I think, I think that we're going to start to see a lot more people train their own sovereign models, their own enterprise models. And so by mid this year, it's important to us that we demonstrate that that is possible with the technology that we're working on. And in a year, a year and a half from now, I'd like to go beyond 100 billion parameter models.

[00:39:15] So, so, so the parameter count of the model is, um, very much the number that we think about when we consider what it takes to make this technology something people take seriously. And I think we need, we need to have a portfolio of models that we can point out and say, these are just as good as the ones that you would have get, you would have gotten in a, in a, in a centralized context, but at 10% of the cost, at 20% of the cost. So I think then people really start to take notice and it becomes an interesting proposition and that's where we'd like to get to. Yeah.

[00:39:44] Uh, and the, um, you, you, we were talking about people stockpiling Mac minis. You're the people that are joining, uh, or contributing, uh, compute to, uh, Macro cosmos are GPU, uh, people with GPUs, uh, spare available GPUs, right? Not CPUs. It's yeah.

[00:40:09] Uh, we're not going to get anything done if we wait for CPUs to do the work, unfortunately, perhaps there are workloads in the future. Um, as I mentioned before, we can imagine this technology is something that allows you to do more than just model training. We we're, we're working on a more fundamental infrastructure problem, not just a model training problem in many respects. And there are a lot of scientific problems, uh, computation expensive scientific problems that are CPU limited. And in those cases, absolutely. It would be very interesting to apply this technology in those cases.

[00:40:39] But, uh, today we support effectively CUDA devices. Uh, and, um, and, um, or silicon devices, which are Mac minis, the new Mac books. Is this online yet? I mean, can, are people training on this network yet? Uh, and if not, when will it be available? Do you have people lining up? Yeah, you, you can find us at iota.microcosmos.ai.

[00:41:04] That's I O T A, which I should probably have introduced is an acronym for the incentivized orchestrated training architecture. Okay. So I O T dot microcosmos.ai is where you can learn more about this project today. We are, we don't have any customers in yet. As I mentioned, we're coming out of research mode right now. We're trying to really calibrate the system and tighten this thing up.

[00:41:27] There's a lot of moving parts that we want to make sure we have really nailed down, but we have startups that are already, uh, ready and committed to working with us training models towards the second half of this year, which is incredibly exciting. So we'll have some collaboration results and some early partnership results by the end of summer. Okay, great. And, and, and, and one other question, this is for compute, but we, we spoke about, uh, federated, uh, computing.

[00:41:54] Uh, could you register data, uh, on the network or, or, uh, make data available and then, uh, uh, for, for pre-training and then have, uh, uh, uh, macrocosmos, uh, send the model around to different, uh, blocks of data, uh, to train the model. It's a very interesting idea.

[00:42:22] One of our other projects in BitTensor is actually a web scale data scraping service where we decentralize the, uh, the efforts of, of hundreds and hundreds of miners that scrape, um, social media data for us. And we can use that for many reasons, anything from journalism to marketing, brand analysis, all the way to AI model training. So we actually have an entirely standalone project, which is dedicated to data scraping, which we think is an important part of our, um, uh, I call it sort of virtuous cycle.

[00:42:52] We have data, we have compute. We also have a lot of, uh, what we call innovation networks, which I didn't get much time to talk about today, but to your, to your question directly about whether we could actually outsource, uh, client side data and use that for model training. My answer is we already are. Uh, we actually just have it in a dedicated, uh, highly scalable system called data universe. Okay. Okay. Uh, well, this is fascinating.

[00:43:18] I'm gonna be paying attention, uh, to, to see how those scales. Do you think this model will be replicated by others? Uh, if, if you, yeah, I, I think it's very important that it is. Do you, do you mean the, um, the, the, the technology or the actual downstream model artifact, the technology, the technology, because as you said,

[00:43:45] uh, the, the appetite for, uh, compute for, uh, pre-training compute is going to, uh, grow the, you know, the number of models out there is only going to grow. Uh, and compute is expensive and a lot of it's underutilized. Uh, so it just feels like, uh, like there would be other people doing this.

[00:44:14] Uh, and, and another question is, are you working with any of the big cloud providers, uh, to, to maximize their, uh, GPU, uh, utilization? So to your first question, I, I certainly, I see other teams right now that are working on very similar projects. Some of them are focused specifically on model training or others are more oriented towards this orchestration technology.

[00:44:44] I think we're one of the very few that's combining those two in the way that we're doing it right now. So, um, I do expect the world of devices, you can call it the internet of things, you can call it whatever you want. But I think that there is a second law of thermodynamics at play here, which is things get more connected over time. And I think that especially when there is a lot of value added and a lot of potential to having multi, uh, the orchestration of multiple devices to do higher order workloads.

[00:45:14] And in our case, there's simply not enough memory on one Mac mini to train the models that we want, which is why we have to do all this work to glue them together. And there's many, many other problems where there's just not enough resources available on one device to do what it is that you need. So I think that there's a lot of other industries that are already thinking carefully about how can we get access to the kind of compute in that specific topology we need it. And I think that what we're building is going to be very interesting to them. Um, devices, as I said before, devices are only going to get more intelligent and more autonomous.

[00:45:43] And they're going to be become much more useful if you can stick them together, like Lego bricks and make these bigger pieces out of them, these different constellations. So I see a bright future for that effort right there, whether we are the ones that really take this all the way to the top of the mountain, or we get halfway there. Um, I'm proud of the work that we're doing. And I think that it's, again, I think it's very valuable, um, for everybody that we work on this. Okay, uh, great.

[00:46:08] So, um, Iota dot macro cosmos dot dot AI. Yeah, you can find everything you need about us just at macro cosmos dot AI. And you can find all of our projects and our research and decentralized AI. And you can find all of our projects and our research and decentralized AI. Bye. Thank you.