
Sovereign AI in Poland: Language Adaptation, Local Control & Cost Advantages with Marek Kozlowski
December 6, 20251h 29m · 15,494 words
Show notes
Marek Kozlowski, Head of the AI Lab at Poland's National Information Processing Institute, discusses project PLLuM (Polish Large Language Models). PSA for AI builders: Interested in alignment, governance, or AI safety? Learn more about the MATS Summer 2026 Fellowship and submit your name to be notified when applications open: He shares how countries like Poland can achieve AI sovereignty by training small, locally-adapted models for specific languages and…
Highlighted moments
their quality in the ability in knowledge about the Polish language and cultures going down. Because, for example, they decided to be focused mostly more on the software developer assistance.
“you have to at least have, after the duplication and filtering out, you have to reach at least around 10 billion tokens. If you don't have 10 billion tokens, it's not worth to perform domain adaptation.”
Transcript
Introduction
0:00This podcast is sponsored by Google. Hey folks, I'm Ammar, Product and Design Lead at Google DeepMind.
Introduction
0:01We just launched a revamped Vibe Coding Experience in AI Studio that lets you mix and match AI capabilities to turn your ideas into reality faster than ever. Just describe your app and Gemini will automatically wire up the right models and APIs for you. And if you need a spark, hit I'm Feeling Lucky and we'll help you get started. Head to ai.studio slash build to create your first app. Hello, and welcome back to the Cognitive Revolution. While we often discuss sovereign AI in the Silicon Valley AI bubble,
0:31we rarely hear directly from the technical leaders who are actually leading national AI projects. And so today, I'm very glad to share my conversation with Merrick Kozlowski, who's leading Project PLUM, which stands for Polish Large Language Models, in his role as head of the AI lab at the National Information Processing Institute of Poland. Poland with a population of 38 million and GDP of roughly 1 trillion, roughly 10% and 3% of the United States, respectively, is an interesting and in some ways a representative case study.
1:03It clearly doesn't have the resources required to compete with the U.S. and China at the AI frontier, but it does have strong technical talent, a real sense of pride in its language and culture, and a deep desire to control its own technological destiny and avoid domination by global superpowers. So, what does that mean in practice? As you'll hear, Merrick's strategy relies on the core belief that by training small models for a particular local language and cultural context, countries like Poland and projects like PLUM can compete with
1:36the latest frontier models, all while retaining control, preserving data privacy, and achieving a major cost advantage. In this conversation, we dig into the strategic realities that motivate projects like PLUM and the technical challenges that they have to overcome to succeed, including how today's frontier models, which are trained on overwhelmingly English and Chinese data, fall short in other languages, why this problem is actually getting worse from one generation to the next as frontier model developers prioritize things like coding performance above support for niche languages, how EU regulation prevents European AI builders from conducting
2:11massive web scrapes and instead forces them to rely on more focused data curation projects, how the Polish government is thinking about investing its finite resources across data, compute, and talent, the language adaptation techniques that Merrick's team layers on top of Lama and Mistral base models so as to inject local knowledge without needing to start from scratch, why they haven't yet had to worry about developing a constitution or other explicit articulation of values for Polish AI systems, and why government agencies and national champion companies are often better served by smaller models fine-tuned for specific tasks and
2:48served locally than by massive generalist models served from the cloud. Overall, Merrick's mix of realism about the challenges of competing with global leaders and his positive vision for transparently created, locally controlled AI is a great window into what AI leaders around the world are thinking and doing to maintain AI sovereignty. So with that, I hope you enjoy this deep dive into the meaning and training of Polish AI with Merrick Kozlowski. Merrick Kozlowski, head of the AI Lab at the National Information Processing Institute of Poland.
3:20Welcome to the Cognitive Revolution. Welcome, everyone. I'm excited for this conversation, too. We met not too long ago at an AI event in Las Vegas, the Enterprise Technology Leadership Summit, and I thought it was really interesting to double-click on everything that you're doing because in the United States and in the, you know, sort of Silicon Valley AI circles that I spend most of my time in, there is this ongoing conversation about sovereign AI. And I think it's funny that a lot of this conversation happens in the Silicon Valley
3:54bubble and sort of makes a bunch of assumptions about what other countries feel the need to have, you know, aspire to create, you know, what's driving those decisions. And I don't hear too much from primary sources of people that are actually doing the sovereign AI projects around the world. So I was excited to meet you and learn more about what it is that you're doing in Poland. Poland, obviously, I think, obvious to me, you know, is a country with a lot of technical skill and, you
4:26know, very distinct culture, obviously its own language, proud tradition, and so I'm really interested to get into it and figure out what sovereign AI means in the context of Poland. Once again, thank you for the introduction and for introducing my person and showing the idea. The idea is that I called it slightly broader, not only the sovereignty, but also the creating the localized elements, yeah, because the localized, it can be the national elements, but also the domain-oriented elements, yeah, and I create the,
4:58maybe not I create the idea, but I am promoting the idea of the localized elements. It means the elements adapted to the language or domain because they can be also a domain, and they are in this domain or the language, they have higher quality understanding, text in this language or domain and have the higher quality, and they are able to create the higher quality text in a generation step. It means that building the localized LLMs, they can be, of course, adapted to the language or domain, has two
5:32goals. First of all, to improve the understanding in this domain of the language, but also give the possibility to generate the higher quality texts, of course, in aspects like the linguistic and cultural aspects. They are the idea, and our goal is to create the models that are in an order of magnitude smaller than the closed, popular now LLMs, but in the aspects of the language and the cultural or the domain, they have the same quality as the 10 times bigger models, and they are open source, transparent, secure, and as much
6:07organic as we can. Yeah, okay. Great. That's a great start. Can we maybe take one step back if we can and just talk about like why this is needed, first of all, from a capabilities perspective. Famously, I think it was, gosh, it's been a minute, but I think it was the GPT Instruct series. I think the model originally was Text DaVinci 002, if I recall correctly, one of the first models that OpenAI trained to follow instructions, they reported basically, we just trained this thing to follow instructions
6:41in English, and lo and behold, it seemed to be able to follow instructions in other languages too, which was, you know, obviously a strong example of emergent capabilities and transfer learning, you know, positive generalization, all these sorts of phenomena that had been, you know, kind of elusive, but I think in many ways like characterize the phase change that we've gone through from earlier AI systems to these more general AI systems, positive transfer being obviously a huge one. But, you know, that's where they started in terms of like just English and, you know, oh my God,
7:17it works in other languages. Since then, of course, they've gone and done a lot of work to try to collect data in other languages to try to even things out. And my sense from just kind of benchmark data is that they have made pretty good progress, but still like performance is best in English. And then you can kind of think of like performance getting worse, sort of the farther a language is from English in the language tree and also just the, you know, correspondingly how many resources
7:51it has, right? Few low resource languages are obviously going to be a bigger challenge than higher resource languages. That's my sense of like... Yeah, at least because I have to correspond to your insights. First of all, the 90% of the training data is English and Chimes. Yeah, when you, even if you look at the biggest open source or the biggest closed LMS, 90% plus data are English and Chinese ones. Only 10% or less are the other languages. For example, it varied, but for example, in some models, the Polish language, I use this example, there is about
8:281% of the corpora or even smaller. And this decides that the vast majority of the skills and the competencies are gained by the English and Chinese instructions. And of course, if you have the large model, it has a huge competencies in transfer learning. You can say they extrapolate. You can very easily extrapolate between the tasks. But for example, even if I have the instructions, I have a lot of mathematical calculations prompted or commanded by the English language.
8:59and even if I ask to do it in the Spanish, the very large model can do it in the intermediate steps. It can translate on the fly the commands from Spanish to English and somehow map the knowledge from the English to resolve the solutions even if it was not learned in the Spanish examples of how to calculate some mathematical formulas. But what is the most important is that this works very good, but this works the same
9:29way as we or our kids learn the language. First of all, we learn how to understand and hear, listen. Next, how to write and how to speak. And of course, if you learn a new language and we get some command in our minds, we try to map these commands to what we know from the primary language, our native language. And the same is going inside the NLMs. For example, I can give you the example that the models that were not trained by the huge volume of
10:04Polish texts, they are still able to be communicative and create the text that is understandable. But there are some statements on some phrases that are very easily identified that they are not the natives. For example, I give you the example about writing the emails in Polish. And for example, they have such a formula typical for English. I hope you stay in good health condition. It's typical for English, but not typical for Polish. And even if you translate this word to word, it's communicative, understandable, but not typical for our language and culture.
10:40So is there more to say about how the leading commercial models are underserving the Polish market than that? I mean, I have the sense that there is a little bit more to it than just the cultural idiosyncrasy. Because even when I look at an MMLU benchmark, it does seem like performance degrades across the language spectrum, right? Like the highest MMLU score is in English. It does seem to get worse in other... I know, but for example, when you look at the benchmarks, we live in the words that we are biased
11:17by the benchmarks. It means, for example, the MMLU benchmark is mostly you choose the solutions A, B, C, D, the multiple-choice questions. They are not testing the ability to communicate fluently in the language. Most of the benchmarks don't test how good is the model in producing the longer forms, longer writings or longer sentences. We usually test the understanding, extractive competencies, summarizing competencies and many knowledge about the facts in the world, but there are, I know, a little or almost a little, very little, very few benchmarks that test how good the
11:53model is in generating the longer forms of text in the other languages than English and Chinese, for example, the niche languages. Because it's much harder. For example, we in Poland create the benchmark PLCC, Polish, linguistic and cultural competency benchmark. And this benchmark enables us to evaluate how good is the model in different subcategories. And for example, there are not only categories of the grammar vocabulary, but also about our culture, tradition and history and many, many others. But we would like to not only evaluate how good is the wordings of the model,
12:30how good the model is in some typical for our tradition wordings and the phrases about the history also, but we also try to check how good the models are in the general spectrum of the using some ambiguous words in Poland. But there is still, but there is still not, this benchmark still doesn't validate how good the longer sentences are in Polish, how good the model is in producing the longer structures in Polish language. Yeah, interesting. So is it fair to say that the primary focus of your work in creating Polish native models
13:07is on these sort of softer skills? It doesn't sound like you're focused on closing the benchmark gap or the sort of reasoning gap that exists between English and Polish. It's more about, as you said, like culture values, tradition, history, cultural competence. But it's because I think that the language is not only the wordings. The models can have the very broad vocabulary, but they should be able to use it properly in the context. And sometimes the language is not only the words. They are the culture, tradition, history.
13:43Everything is mixed in it. And in order to create the model that behaves as natives, you have to inject not only the knowledge about how to create grammatically correct sentences, but also how to use special idioms or phrases in special context. Or what places are typical for Polish history or maybe what places are viral now. Generally, it's like you have to mix the history, grammar, vocabulary, art, entertainment, culture, and tradition, everything into the mix
14:13to create the language ability that is somehow similar to the natives. But ask me why we are doing that. First of all, because as I mentioned sometimes that we believe in the idea of the localization elements, the localized elements, adapted to the language, that there are as much similar natives as possible. The second issue that the competency gap, for example, we believe that we have to develop our people or our engineers to have the skills to build our own models, because maybe in a few years the market
14:50will change, maybe the mobile will be closed, or maybe some models will be forbidden. there are plenty of models currently in the European Union that we are not able to use because of the IAC. Even in the licenses, the LAMA 3.4, the KIMI and many other models, they have the sentence, the statement, in their license that they are prohibited to use in the European Union. Maybe we will be forced to use this knowledge to beat our own models.
15:22Maybe they will be a little bit worse than the Chinese or USA, but they will be our own. sometimes it's better to have the competencies to beat even something a little bit worse, but they have the ability to do it, then don't have this. We can. Sometimes it means more than you think. But also we think that in this approach, in the Plume family, because we create the family of models, we also believe in the transparency because we show how we built it from the scratch.
15:57We released a few weeks ago or two weeks ago, sorry, we released the publication of almost 100 pages how we built these models. And also, we not only released the publication, the recipe book, the cookbook, but also we published on the honey phase the samples of our data sets, the instructions, preferences, because we would like to show not only the open weights, but more more than because the open cells are not only the open weights.
16:28There are also some samples of open data and the cookbook how we do it step by step in a very detailed manner. And what is for us important, because I think even now they are the most popular open source models are the Chinese now, but they are only open weight. There are no samples of the instructions or preferences they use to train the models. And we would like to go a step farther, to be as transparent as possible. And also, we invest lots of in their organic data, because we believe that, we also prove that, that when you have three stages when you learn the models. First, the pre-training. It's somehow similar to learning the kids a new language, that you identify the words, how to create structures from
17:03these words, and some pieces of information. But children after this type of learning is able to resolve the mathematical calculations or write the essay. It's like you learn the language, but you don't learn the competencies. Next stage is the SFT, supervised fine-tuning, you learn how to resolve some tasks, downstream tasks, write the poem, write the essay, summarize this article, perform some calculations, you learn the competencies. Like the children in the school, you have the math, the geography, chemistry, and many others. And after all, you have the alignment, the preference learning, that you mark what the children have done during the test, for example, and this information, these marks, makes them what should be corrected or not. And with the same what we have done with the kids, that we learn the language, the piece of information,
17:36the wordings, and the structure, next we learn the competencies, and we evaluate them, and during the feedback loop, we try to improve their abilities, the same things are done with the RLMs. And for example, when you are doing the pre-training, you show the model the hundreds of billions of tokens to learn the language, and after all, in the SFT, the supervised fine-tuning stage, we show the model the syntactic instruction. Syntactic means that they are produced by the other RLMs. Usually, if they are linguistically poor, they also degradate the model. because in any stage of the learning, when the models see the poorer data, it will degradate, it means it's going down, there is quality of the linguistic creation of the sentences, and so we focus mainly on creating the
18:09organic data sets, organic instructions and preferences, or even if you use the RLMs to produce such instructions, we check it by the humans to improve their structure and quality. And I think there are the novelty. The second novelty is that first of all, there is the open source, open data, and open cookbook. The second is transparent, because it's written in the cookbook what we have done step by step, and we show the samples. We also focus on the organic data, organic instructions, organic differences, and I think this is the reason why the GPTs and other people are so good, because they are also they have plenty of manual instructions, and they don't show them because there is an intellectual property of these companies.
18:39And we also, we secured our models on our own, because we discovered that these models are secured for the English speakers, they can be much easier hacked than secured for the Polish speakers. I think they are novelties, maybe briefly speaking. Hey, we'll continue our interview in a moment after a word from our sponsors. If you're finding value in the cognitive revolution, I think you'd also enjoy Agents of Scale, a new podcast about AI transformation hosted by Zapier CEO Wade Foster. Each episode features a candid conversation with a C-suite leader, from companies including Intercom, Replit, Superhuman, Airtable, and Box, who's leading AI across their organization, turning early experiments into lasting change. We recently cross-posted an episode that Wade did with OneMind founder and CEO Amanda Calo about AI-led sales.
19:11And I also particularly enjoyed his conversation with John Nerona, chief product officer of AI Product Pioneer and recently minted double unicorn Gamma. From mindset shifts to automation breakthroughs, Agents of Scale tells the stories behind the enterprise AI wave. Subscribe to Agents of Scale wherever you get your podcasts. Are you still jumping between multiple tools just to update your website? Framer unifies design, content management, and publishing on one canvas. No handoffs, no hassle, just everything you need to design and publish in one place. Framer already built the fastest way to publish beautiful, production-ready websites, and it's now redefining how we design for the web. With the recent launch of Design Pages, a free canvas-based design tool, Framer is more than a site builder.
19:42It's a true all-in-one design platform, from social assets, to campaign visuals, to vectors and icons, all the way to a live site. Framer is where ideas go live, start to finish. And now, they've added a Framer AI layer to make it all faster and easier than ever. With Wireframer, you can skip the blank canvas and get a responsive page with structure and starter content ready to edit. With Workshop, you can create new visual effects, cookie banners, tabs, and more. No coding needed. And with AI plugins, you can connect top models from OpenAI, Anthropic, and Google to generate images, rewrite text, generate alt text, and more. Ready to design, iterate, and publish all in one tool? Start creating for free at framer.com slash design and use code cognitive for a free month of Framer Pro.
20:15That's framer.com slash design. Use promo code cognitive. Framer.com slash design. Promo code cognitive. Rules and restrictions may apply. I have like seven follow-up questions I want to ask about various parts of that. And maybe we can kind of break it down by inputs to AI for one thing, right? Obviously, the big inputs are data, compute, and talent. And you touched on certainly data and talent there. I also do want to come back to the safety training, because that's always a keen interest of mine. But maybe let's start with like the goal. You've spoken about it somewhat, but I think one big challenge that
20:51we have, certainly in the United States, and we have all this talk, right, of, especially in the context of the geopolitical competition in AI, there's a lot of talk about, well, we want to have AI with democratic values when we don't want to have Chinese values, or maybe we even are bold enough to say we want American values to be the values that the AIs embody and kind of propagate through the world. That obviously brings a big question, which is, well, what are those American values?
21:25And I can certainly say that there's no like single agreed-upon answer for that, right? What American values are is like hotly contested on an ongoing basis, and that leaves basically the AI companies to try to come up with their own, you know, best guess of what that should be, and, you know, that too is like often sharply criticized because, you know, it's too woke, or it's, you know, not woke enough, or it's, you know, right-wing extreme, or it's, you know, it's describing itself as Hitler in some of the LLMs, they
22:00are somehow they are the compressed representation of what we have in our web in the internet, yeah? They are somehow, what topics are the most important, what information is the most popular ones, it somehow is reflected by the LLMs, yeah? If you have the problems, the political problems, the religious problems, the everything, what is there is also reflected somehow in the compressed LMs, because LLMs, they are somehow, they are compressed memory repositories, yeah? They are the compressed stores of the memory of the internet.
22:34Certainly that's, you know, all that stuff is baked in. Sometimes, I don't know how far American, the leaving companies have come today in terms of filtering the training data. I know that there are some techniques that are like, we're going to get rid of all the bad pre-training data, you know, and just try to show this. Yeah, there is a typical step, yeah? Even in our project, there is called data curation, that even because, as I mentioned, in the pertaining stage, the
23:0690% of the data, the web data, and the web data, plenty of them is creepy, yeah? You are not able to use them, because the model will be not stable. And in this data curation step, there are two sub-stages. The duplications, you remove the same information that is repeated very often in the internet. Sometimes it's scaled to two times, because there is plenty of duplicates in the internet. And also there is the filtering out. We filter out the data that are very crappy, it means they are
23:40low quality. It is, for example, there are plenty of special characters, plenty of interpunctions, plenty of not being recorded in our vocabulary. there are plenty of such disturbed data, which should be truncated, because it will have impact on the stability of the quality of models. And I think you mentioned that the big companies, they have such tools that are not even able to eliminate some poor quality data, but they even eliminate some, for example, the theories,
24:11some points of use, and much much broader section, not only the linguistic aspects of the data. So it's the same censorship. The Chinese models, if you ask about what has happened in Tiananmen, square, they are not able to give you any information. Generally, the people who build models, they can able to isolate, or how it's called, block the bank, some important information, that for people who are not aware about that, will be the reflection of the world, we told some part of it.
24:41Yeah. But there's at least, like, two layers to this, right? I mean, there is all this data pre-trained, or pre- filtering. And I genuinely, I could certainly believe that the Chinese models are trained on data that's so thoroughly filtered as to never have seen, you know, any document about Tiananmen. I think that even not about in the pre-training stage, I think in the last stage, because as I mentioned, there are three stages during the learning models. The pre-training, the SFP, supervised fine tuning, and the preference learning, sometimes
25:13called the reinforcement learning with the human feedback, but there are other methods like DPO or ORPO. And I think in this stage, they secure their model how to not behave. Right. So that's what I want to get at for what you're doing in the Polish context, because in, I don't know what, you know, the Chinese companies are doing. I do know that the American companies are developing their model specs or their constitution. You know, it's basically this super long document that says, this is how we want our AI to behave.
25:45And to their credit, you know, they're starting to be reasonably transparent about what those are, so that at least the public has a sense of like what they're going for. But again, it's like, it's in the US context, it's like pretty contentious because, you know, everything is contested here. In the Polish context, is it like that? Or is it, you know, is it an easier time? Like, do you have a constitution for what you want Polish AI to be? There are some strategies. There are the strategy, how the AI should behave, or no, maybe not, how it should
26:21not behave. Maybe in this case, for example, there should be ethical, that should not blame anyone, or for example, to avoid some topics that are, for example, very risky topics, rather, how to say, the topics about the hate speech. Yeah, there are some places that are typical, the same for the models from the Chinese or the USA, they are the place when there is a risk that the model behaves in an unethical way, or in a way that we can be blamed, that it's not the, it's not, it's rude, at least, at least rude.
26:55But, of course, there we go, some political tensions, yeah? But generally, I think we don't have such a huge, now currently, very huge constraints, as you mentioned, yeah? We don't have a constitution that we have plenty of points that you have to obey. I feel we are mostly that the models should be ethical as much as we can, but we don't give too many constraints to the model, because I think we are on the other level of development of the London China, Chinese or the USA government or the companies, because we are, I think, a few years behind them, or
27:31two years or three years, it's hard to say, but generally, I think we are not, we don't have such a, we have our own regulations, but not the regulations containing the, how the model should behave, but rather what kind of data we are able to use for training. We have many constraints, things like that, focus on the data, then how the model should behave. Yeah, interesting. So, I mean, I've never even been to Poland, so obviously, you know, should be very humble in terms of my ability to describe it, but one high-level fact that I know is that the large majority of
28:08Polish people identify as Catholic. So, it seems to be. For many years, it used to be a good statement. Currently, I think it depends on how the city is big, that the inhabitants of the cities versus inhabitants in other villages, I think they should be valued. Yeah. So, how do you think about that dimension? You know, should the, I just happened to have done an episode not long ago about Catholic AI with a company that is literally building AI that embodies Catholic values, you know, specifically for religious Catholics.
28:39But, in your context, you know, you've got this sense that, like, okay, well, you know, maybe a majority of people are Catholic, but maybe that's on the decline, and maybe it depends on, you know, an urban-rural divide. Do you have, is there some sort of decision-making process where you think, okay, like, how Catholic should our Polish AI be? And, you know, does it vary in different situations? Is that, are you guys getting explicit about articulating goals there? I think we are much more liberate, you know. I think there are, now, now we don't have such ideas
29:13to create the reflection of our world. But, as I mentioned, I think there is very, can be very easily, the models can be very easily constrained by the, by the preference learning, and you can learn it to behave in such a special way. But, now, currently, when we produce the family of the models, we produce not only the chat models, but also the instruct models, the base models, we give the possibility to companies to use any kind of the models, because we know that some constraints may have a disruption effect on the, some business cases.
29:46But, generally, I don't think that we, as the producer of the, of the, as the releasers, or the, I don't know, the, the builder of the models, we should get all of them, the people, and we will be, we decides how to skew them. Gotcha. Okay. Yeah, very interesting. Do you, do you envision that this will become something, you know, as, as you presumably go on to train more in future models, and they become even more powerful, and, you know, potentially, I don't know to what degree
30:17you sort of aspire to serve a consumer use case versus, you know, in empowering businesses in the country, but do you think that this becomes a challenge at some point? Do you, do you envision a future where there is a sort of Polish constitution for AI that actually seeks to answer that question? And if not, like, how do you think you ultimately get around that? Because it seems to be a very central thing that the American companies feel they need to grapple with. So, if, if you, if you think you can avoid that problem indefinitely, I'm kind of wondering.
30:51I think we have much harder problems because, for example, you have the AI constitutions in the companies and the USA market, but, for example, you can very easily use all data you have without any constraints. Of course, there is a problem with the, with some suits and many other cases, but there is a, like, maybe it's a long process, not very, I think for most of the companies in USA, they can, they can take this risk because they are still profitable enough to even to pay some, some, to say, some, there's some
31:24bad decisions by the jettison or and so on. I have to say the, the arbitral decisions, but generally, I think in Poland and European Union, we have the AI act and our local, local regulations like, for example, the, the acts concerning the, the, how it's called the rights, the authorship rights. And fair, I think these problems, the legally speaking, these documents, we have much harder impact on our quality of models than, I think, then, then we can, I think it's enough, it's enough huge constraint not to go in
31:55farther. Yeah, because, as I mentioned, in the European Union, we have the AI act concerning the general purpose elements. We have also our local regulators, like, for example, the AI act about the, the act about the authorship rights, and both of them combined create a much harder constraints than, for example, any kind of constitution that is much, I think, flexible, more flexible than our regulations. And I think we don't think currently about the AI constitution, but I know that maybe in that one, two years, something but talking about that a few, but currently in European Union, Poland, we fight with the currently existing and obeying the regulations. Yeah. And I think they are much harder, much harder, and they are much impactful than those you mentioned in the USA, because for example, you have the AI act or the authorship rights regulations.
32:27For example, they can eliminate the 80% of the data from your training, training datasets, and so it's a huge, they have a huge, this has a huge impact on the quality of, of models. Yeah. Everyone listening to this show knows that AI can answer questions, but there's a massive gap between here's how you could do it, and here I did it. Tasklet closes that gap. Tasklet is a general purpose AI agent that connects to your tools and actually does the work. Describe what you want in plain English. Triage support emails and file tickets in linear. Research 50 companies and draft personalized outreach.
Build a live interactive dashboard, pulling from Salesforce and Stripe on the fly. Whatever it is, Tasklet does it. It connects to over 3,000 apps, any API or MCP server, and can even spin up its own computer in the cloud for anything that doesn't have an API. Set up triggers and it runs autonomously. Watching your inbox, monitoring feeds, firing on a schedule, all 24-7, even while
33:00you sleep. Want to see it in action? We set something up just for Cognitive Revolution listeners. Click the link in the show notes and Tasklet will build you a personalized RSS monitor for this show. It will first ask about your interests and then notify you when relevant episodes drop. However you prefer. Email. Text. You choose. It takes just two minutes and then it runs in the background. Of course, that's just a small taste of what an always-on AI agent can do. But I think that once you try it, you'll start imagining a lot more.
Listen to my full interview with Tasklet founder and CEO Andrew Lee. Try Tasklet for free at Tasklet.ai and use code COGREV for 50% off your first month. The activation link is in the show notes, so give it a try at Tasklet.ai. Being an entrepreneur, I can say from personal experience, can be an intimidating and at times lonely experience. There are so many jobs to be done and often nobody to turn to when things go wrong.
33:33That's just one of many reasons that founders absolutely must choose their technology platforms carefully. Pick the right one and the technology can play important roles for you. Pick the wrong one and you might find yourself fighting fires alone. In the e-commerce space, of course, there's never been a better platform than Shopify. Shopify is the commerce platform behind millions of businesses around the world and 10% of all e-commerce in the United States. From household names like Mattel and Gymshark to brands just getting started.
With hundreds of ready-to-use templates, Shopify helps you build a beautiful online store to match your brand's style, just as if you had your own design studio. With helpful AI tools that write product descriptions, page headlines, and even enhance your product photography, it's like you have your own content team. And with the ability to easily create email and social media campaigns, you can reach your customers wherever they're scrolling or strolling, just as if you had a full marketing department behind you.
34:05Best yet, Shopify is your commerce expert with world-class expertise in everything from managing inventory to international shipping to processing returns and beyond. If you're ready to sell, you're ready for Shopify. Turn your big business idea into cha-ching with Shopify on your side. Sign up for your $1 per month trial and start selling today at shopify.com slash cognitive. Visit shopify.com slash cognitive. Once more, that's shopify.com slash cognitive. Okay. That's interesting. Hey, we'll continue our interview in a moment after a word from our sponsors. So turning to data then, we can check back in on the state of the Polish AI constitution in a year.
On the data front, you had mentioned that in the biggest open source models, maybe 1% of the data is Polish, you know, quick back of the envelope math, I think they'll, you know, the Lama models are like maybe up to 15 trillion tokens that they've been trained on. I don't know if they disclose their data mix, but you know, that would cash out to something roughly on
34:38the order of a hundred billion tokens in Polish that the biggest projects might be using. I understand you have quite a bit more data than that, but also your last comment. We don't have the trillion of tokens. Yeah. Because even as I mentioned, even in the Lama, there was about, as I mentioned, the one person who's the Polish language or maybe less, we have now the several hundred billion tokens, you know, we don't have even the trillion of tokens because as I mentioned the duplication stage, the duplication stage and the filtering out stage, they eliminate lots of data.
And we don't have one trillion tokens after these data curation steps. So where are you getting your data? And your comment about like, um, the, the difference, the sort of regulatory arbitrage that the American companies are potentially taking advantage of, are they able to use some Polish data that is on the internet? But no, no, because remember, I remember that it was some times ago that there was the
35:11people who analyze the crawlers that are going from the, on the webs, on the websites in the Polish, in Poland, they identified, there are plenty of anthropic crawlers and there was plenty of robotics pages on the websites that disallow these anthropic crawlers to get data.
I mean, I mean, there are plenty of the crawlers from the US companies, researching companies that are crawling the Polish data, even if they are not, if they are, if, even if they are not allowed to do it, because as I mentioned, it's much harder, for example, to be the, I know, they have the, the, the, the court in the USA and to have some, I accuse them for using the data and able to, to fight with them in the, on the USA court.
Even if, even if, even if they have some proofs that they use data, that they, that they, that they, they have the desire clauses. And so in that way, you're, if there's a hundred billion tokens that they're getting
35:43off the internet, it sounds like you can only use a fraction of that. And then you had to go elsewhere to find the hundred, a few hundred billion tokens that you've. Yeah. So I'm really using the, the central libraries with the, some, some sources of data that are not web. Yeah. Because as I mentioned, then the vast majority of data used by the, the big vendors also by us, my web data. But for example, you have also some data that are not published in the web and we can use to somehow, of course, to some extent use them. But as I mentioned, there's a lot, there's a, there's a minority of data and still the, even for us,
36:14even if you have some access to the, the local organizations and so on, the, the vast majority of data we use, there is a web data. And the problem is the same for all other players. Yeah. And maybe we can much easier identify some, some web sites that are not easily crawled by the external crawlers. But generally I think the, most of the companies, the open AI or the anthropic, they have the, I don't know, maybe 80 or 90% of our data still. Hmm. Yeah. So what, where else are you going to get data?
36:47Like what is your data process? And we have, for example, the first of the massive mesh of data, but also we have, for example, the, the, the, there is a called the, how is it called the, uh, the library of the science. Yeah. The, the, the, there is a plenty of, uh, for publications and so on. We also have done some private, bilateral agreements with the, with the publishers, but not published on the web. But as I mentioned, there's only the fraction of data we have in our corporate.
37:19And are you also, you mentioned kind of doing a lot of human review. How is there a, the huge, the, the, the, the, what we have, our advantage is not in the data used for the pre-training stage. Because as I mentioned, I think the 80% of them is still in the, our anthropic or open area repositories, but that we have, uh, uh, dozens or even the hundreds of annotators that create the manual, the instruction preferences. Yeah. And because they give us the ability to create the new data that is
37:51not published in the, on the internet still. Is there like a Polish equivalent of like scale AI or label box that you're working with to do this, or is this a project? No, we have made our, our, our internal tools, not, not, not, uh, not so crowdsourcing for once. So you guys have built your own platform for human preference data. Yeah. Okay. Preferences and the, and the human instructions are built locally, internally. Of course, some of them, some samples of them we publish to show what is the structure of our instruction
38:23preferences and some examples. But most of them is still the closed asset. So how do you think about that? I mean, what, one question that I've been thinking about in the context of this whole sovereign AI discourse is obviously, you know, as a national government, you can have different strategies or different, different goals and different strategies for what you're trying to do. One goal is, as you alluded to, we want to make sure we have our own base, you know, our own data, our own talent base compute. We'll get to in a minute.
38:56Um, so that if we get cut off or, you know, who knows what might happen, we have some sovereignty over, you know, what's going on, right? I think that's some, some possibilities to, to develop in a different way. Yeah. But also there is the second, because this is the one, right? Yeah. As I mentioned, we can something, we can mean something more than you think. Yeah. That if you can do something, even a little bit worse, there is still the, the competency and possibility to do a new ways and new movements. But I think there's also the second, the second issue that I
39:33believe that the AI-agentic revolution will be based on the smart localized models. Why? Because first of all, there are some branches or sectors in the economy or even in the, in our public sector, but we are not allowed to use the cloud-based solutions. And there are some regulations or the, we are not, the risk is too high. And there are some demands to have the on-premise, on-premise models. When you have the on-premise models, you always have some challenges, like for example, the GPUs you have to buy,
40:04the energy consumptions. And usually when you realize that, for example, you need to buy the 16 GPUs and the, and the pace for energy, it always goes to the downscaling. Yeah. To, to use the, as small model as possible to achieve the expected goal. And from our experience that, for example, the people, and especially the businesses, but also the, the public sector, they don't demand exactly the chat GPT. The general purpose LLM that is able to resolve 1000 tasks. Usually in the business and the public sector, we have the demands for the 10 or 20 use cases.
40:38And you are able to create the smaller models that are able to resolve these tasks from the same level as the few short use big LLMs and hosting them on the previous solutions. And I think when you go to the AI-agentic solutions, there are plenty of agents, it means the plenty of models that are used to resolve some complex scenarios. You will have to downscale the models to using only the, as small as possible models to be energy effective. And also the, the, the economy aspects are, are crucial now.
41:11And I think this is the place where the small localized models are, can play well. Yeah. That makes a lot of sense in the business context. And yeah, I, you know, in my experience, I would say the same has been true. Like when I'm really trying to dial in performance for a particular use case, and that's all I care about. And I kind of know that this model is going to be deployed in a controlled environment where, you know, because of the way the system is set up, I know what the inputs are going to be.
41:45I know what the outputs are going to be. I know that, you know, I have a, I have other layers of control. Then yeah, I can just kind of dial into one task or a few tasks. And often a cheaper model, you know, with the right training can do just as well. Especially when you have the, because mostly the people are now using the cloud LLMs in a few-shot manner. Yeah, because they are very powerful, they are able to resolve a very broad number of tasks, thousands of tasks,
42:17and they use them as the, you know, zero or few-shot scenario. It means they are integrating the API of their own systems with the cloud-based LLMs, and they use them as they can out of the box easily. Because you could create the prompt and use the output, that's all. But when you have the, when you need to create the much more controllable solution, the closed solution, the on-premise solutions, you are not able to use some cloud solutions. You have to go with the, with the different, as I mentioned, the aspect of the decisions.
42:50Do I need the multimodular or textual only model? Do I have a training data set? Maybe if I have a training data set, I can, I can supervise, fine tune the smaller model and achieve the same quality as the few-shot applied cloud-based solutions. And mostly when I, we can, we create many deployments currently. We identify that when you have the few or 10, 10 different use cases, and you create for them at least 1,000 instructions or 3,000 instructions, and you SFT, supervise, fine tune the smaller models, you achieve almost the same
43:23quality or maybe sometimes higher quality than using in a Q-shot approach, the big, very big, large cloud-based LLMs. But of course, you have to prepare training data sets, at least 1,000 instructions, the best, the higher number of restrictions is better. But 1,000 is enough. You mostly organic ones, but, or maybe the semi-automatically created with the human factor. And then you have 1,000 or more of the instructions, you can SFT smaller models. And in this one task or a few tasks, you will have the same quality as using the zero-few-shot cloud-based
43:56SEMS. And I think this is the, the future. Because I think if they, if they will go to the business and the business will calculate the risks, the money, the possibilities, the, how they control the solutions, how, what is the impact of their decisions? They will finally, in this AI-agentic environment, choose the small local SEMS models and fine tune them to their demands. And there is also one risk. I will show you that because we discussed it last time, lastly, with the SEMS and of my colleagues.
44:28For example, in the Atropics models, the CLOT, for example, the LLMs from the Atropics, their, for example, their quality in the ability in knowledge about the Polish language and cultures going down. Because, for example, they decided to be focused mostly more on the software developer assistance. And if they focus on the, this might be a kind of more point, a fraction of the market, they go with the, down with the qualities in the, for example, generating texts in the niche languages.
44:58And for example, imagine that, for example, in Poland, you apply such a model from Atropics, you integrate it with your environment, with your ecosystem. And after the next months or next years, the next releases of this model going down on your competencies that you demand. And you have to choose the other model or you roll back if you can, because unless you are not able to roll back, revert the previous models, or you have to choose another vendor. As I think creating the huge integrations based on the cloud LMs is a huge risk because, as
45:34I mentioned, during the next years, they can change their target objectives and don't need it to focus on the Polish or Czech Republic. Or Czech Republic languages, because they are not the market for them. Yeah, that's really interesting. I've never heard that before. So just to make sure I understood correctly, you are seeing worsening performance over time in Polish, on like top issues of sort of Polish culture, general world knowledge, as the cloud models have progressed through generations. Yeah, because we have the, I don't, I can say there's the anthropic based
46:08models, I think the cloud and then the Haiku, or maybe I think the Haiku, but also we identify this problem with the GPT models. For example, the GPT models that don't going up with the quality of the Polish language and culture competencies. There's some of them even going down with the next releases. I think the problem is much more broader. I think that there are, the creators of the models, they analyze the market. And for example, if they need to focus on some competencies, and they improve these competencies, there is
46:42a trade off that other competencies are going down. In this case, in this case, in this case, the Polish culture and linguistic competencies. I can check it, it was the cloud model for Anthropic, and this version is going down on this, our PSC benchmark, when you compare to the previous releases. Wow. Okay, that's a really interesting data point. Yeah, and I guess it maybe sort of answers the next question that I had for you, which is, my general working model has been that the AI frontier model developers want
47:15as much data as they can get, you know, and if you had any data for them, they would be happy to take it and maybe even pay you for it. But what you're saying sort of suggests, well, maybe not always, because, you know, they're trying to do the smallest models that they can as well, while, you know, they're doing all this sort of distillation, they're trying to, you know, go for efficiency, they're trying to obviously serve the core use cases that they're getting paid for, which is,
47:45you know, a lot of coding. And so maybe if you showed up at their doorstep with a few hundred billion tokens worth of Polish data and said, hey, would you like to use this, maybe they, at this point might say not really, because we aren't that focused on that use case versus we'd rather, you know, go do another, however many billion, you know, generations of token of coding tasks and, and use those tokens instead. I guess, would you guys ever consider, I know you have some, it sounds like you have some open data,
48:16but not all of it is, is open. If I was the, if I'm thinking as the government, I guess another goal that I might have is, I want my users, retail users, you know, just my general public to be as well served by AI as possible. And I don't know if you have stats on like what the Polish, you know, retail consumer is using right now. Are they going to ChatGPT? Are they going to Gemini or something else? Mr. All, who knows? You may know, I don't know. But if I was the Polish government and I was saying, okay, here's what
48:50my people are doing. They're using these other companies. We've gone and collected all this data. Is there some sort of deal to be made with the AI companies where you might say, hey, we'll either give you this data or perhaps license you this data. You pay us for it. And that way you can incorporate it into your process. And that way you can serve the Polish market better. It's, I've wondered if there's some trade that. It's a, it's a good point of view that this, I think the next natural step, yeah, that you are not able to get more data without some assist or the cooperation
49:25with other players. But as I mentioned about the previous statement about the, to serve as good as we can for our citizens. And I think there, there is still, there is one of the goals of our projects because about the plume, the family of models that they call plume, Polish language models. It's a family of models, but not only the family of models, but also the assistant and chatbots for citizens and city inhabitants. Yeah. That we do not only focus on the models as itself, because the models are very good asset, but how to build. Based on those models, the chatbots, the RAC approach solutions that can work for
50:00the citizens nationwide, but also in the, for city inhabitants, for the, for the local chatbots, for example, in the cities, the halls. And as I mentioned, it was the, it's very good to mention that, that sometimes we are trying to create better and better models, but the problem is, for example, on the digital level, on that other, another, another place. For example, there is not a problem that the model is a little bit worse or better, but for example, there is no chatbots for the cities, municipal halls and so on. Yeah. And there, there is a, two issues. It's first of all, the deployment issue that this model should be
50:37a little bit customized, supervised, fine-tuned to be able to work as the part of the, of the citizens and, and, and, and, and, and, and assistance for the, for the city inhabitants. And this is the, and the second issue, as I mentioned, that sometimes maybe there's time for some cooperation because the, only the cooperation gives you the ability to improve your data sets and improve your models. I think it's a very good step. And I think as I mentioned, it's, it's, we are, we are developing in our pace, but we know that there is a place where you are not able to go further and
51:11you have to be supported by someone else. It's normal. The same is in the business. To some level of your development, you reach some points, but you have to be supported by some better players or diverse players to be better finally. Yeah. So do you know what that kind of market share breakdown is today? And is there a sort of established goal that you have to win market share with the models or is it? It's very hard. It's about the winning. And I think because the blue models are not the corporate initiative.
51:45It's not the private money and all the funds. This is the project supported and funded by the Ministry of Digital Affairs. It's a consortium of the six institutes and now the eight, because we enlarged it in the second year, institutes and universities. And we are the public initiative. If you are the public initiative, you don't think too much about the investment or the number of customers. We are much more focused on how to be as open as possible, how legal, because you have to be
52:19legal compliance with two regulatory holders, transparent as possible, organic, because it improves the linguistic possibilities and security, and how it can be used by the public sector as much as we can, because mostly in the public sector, the models should be closed on on-premise deployments. Yeah, gotcha. Yeah. You've shared a lot about how you train these models, but what is the base? You're not doing all the pre-training from scratch, right? My understanding is starting with the base.
52:49I can say about it more. I can say it and explain it much more in detail. We're trying to create the models from the scratch. It's from the random weights, but the problem is in the pre-training stage, the number of tokens you have. As I mentioned, even if you look at the APRIL reports from the DSLM models, they prove that even if you have the 8 billion parameters models, you have at least 1 trillion tokens to have the stable training. It means stable that gives you the moderate or high-quality base model.
53:25And in our case, as I mentioned, the duplication stages and filtering out give us around 200 billion tokens. It was too little to create the model from the scratch. And we used, of course, the LAMA. Of course, now the LAMA are much more closed, but one year ago, they were still open and the LAMA license was not so prohibited in the European Union as it is now. But usually we use the LAMA-based models and the MISTRA-based models, and we continue pre-training.
53:57We perform the language adaptation. Language adaptation means we continue pre-training them on our corpora of Polish texts. And after all, we have the new base model. And this new model and new base models can be SFT and difference optimization learning in the next second and third stage. But as you mentioned, we are not able to create the moderate quality or the good enough quality model without 1 trillion tokens, and we don't have 1 trillion tokens in Polish language.
54:28Now we have made some experiments with the mixture of languages. Yeah, we made not only the Polish language, but we also use the other languages, the mixture of language to get this 1 trillion tokens and to make some random pre-training from the scratch. But the efforts will be in a few weeks. Hmm. Okay. On this language adaptation step, I have a couple questions there. One is, do you continue to mix in English? Do you try to preserve the model's ability to speak English?
55:00Or after this language adaptation, does it just only speak Polish? Of course, when you have the cascade learning, because if you use the base model that was already created and you perform the language adaptation, of course, there is always in the cascade learning the problem of forgetting. Some knowledge is forgotten from the previous learning stages. But there is still some persistence, still now. Even if we pre-trained a few epochs on our Polish data, the models still have the competencies, for example, to write something in English. We don't prune any other competencies.
55:33We don't prune any other competencies in other languages. Somehow, they are forgotten because there is a problem of forgetting in the cascade learning. But generally speaking, we don't prune it manually. We only take the base model and perform the language adaptation, continue to pre-trained a few epochs. And this is how we make better the abilities in the Polish language. Of course, there is a trade-off. Some other languages are going down, but they are not pruned at all. And that is also where the world knowledge comes from, right? And by world knowledge, obviously, there is all these sort of local details of life, right?
56:08And you can maybe give better examples. I am sure you can give better examples than what I would give. But I am thinking just like, what are the names of the Polish candies that kids like? And how does one file a document if you want to sell a car to somebody else? There is, you know, surely some filing process. All these sort of little details, they are absorbed in that stage as well, right? Yeah, and you have to know that usually when you, the knowledge, the general knowledge is pertain.
56:39But then you ask me about the factuality, how a model is good in some facts for the regulatory issues, some law issues that are changing in the time. It's always the problems in any kind of LLMs. Yeah, because you pre-train the LLM, for example, on the data that are till March 2025. And you don't have in this memory store information about the changes in the law or some regulations or even some situations, accidents and the names of the new politics after this time point. But generally, we use it when used in Iraq approaches or data.
57:11It means you have the retrieval stage, when you have the database, the knowledge base up to date, it's much easier to be updated. And after all, when you get some retrieval stage, you use the models to synthesize or generate the answer based on that. And in this case, we use this factuality issue, because I think none of the providers of the LLMs, even the big ones, are not able to pre-train it and a new interval of data is coming up. Yeah. Do you think this would work for companies?
57:41This has been sort of a, this is a bit of a digression, but it's been a question I've had in my mind for a long time. Um, like almost two years ago now, I did an episode of the podcast with a company called Mosaic LM. And what they were doing among other things was this sort of continued pre-training for businesses. They would go into a business and say, let's get all your tokens, you know, and that's, this could be all the Google docs that you've got in the Slack history and all these various things.
58:13Let's compile that. And now we can pre-train on that. And hopefully the model will start to speak your, you know, internal native dialect of whatever language you're speaking. I called it to, as you mentioned, the domain adaptation. Yeah. For example, you have some closed data. Yeah. There's a closed, for example, I have my internal closed data for the documents about my insurance, my customers, about some reports that are not open. And I would like to pre-train the model on the, on this data to make it more adapted to the
58:43domain. And we have done such a project now for the biggest bank in the central Eastern Europe. The PKO BP is one of the biggest bank in Europe, the biggest in the central Eastern Europe. And we have performed what you mentioned, the domain adaptation. We adapt the models to their domain. They have some, they have said that their own data, domain data that is closed data. And we pre-train, continue pre-training the models on their data. And I think it's a very good, very good approach. We proved that this, in different tasks, it varied, but there are some tasks that this domain adaptation gives you
59:20a huge gain in the quality and in some, in some financial measures and so on. But I think there is one remark I have to mention. And I think that only the huge companies has enough data to be worth to perform a domain adaptation. Because we know that, for example, you have to at least have, after the duplication and filtering out, you have to reach at least around 10 billion tokens. If you don't have 10 billion tokens, it's not worth to perform
59:50domain adaptation. And I think if you would like to have the 10 billion tokens in a domain Coquera, you have at least 30 billion tokens before that, the duplication and filtering out stage. And the 30 or 40 billion tokens, I think there are only a few companies, maybe not few, but less than 100 in Europe that has 100 billion tokens, they closed internal data. Yeah, I guess that, my intuition is that, I guess it depends what you count, right? I mean, excuse me, I've, I started a company that's, you know, 40 people and I don't know how many
1:00:26tokens we have, but over all the Slack messages, all the Google Docs, all the Jira tickets, all the, you know, contract proposals that we've sent and the revision history on all of those, I do feel like it adds up pretty quick. So I guess maybe one of the barriers is just like exactly how deep are these companies willing to mine into their own data? Like if they're actually willing to go get email data from their employees. I think if you count this data, there are maybe the plenty of billions of conversations, the hundreds or thousands
1:01:02of agreements or the proposal of agreements. But if you sum up them, and they count them, and after all, they duplicate them and filter out, it's very hard to get 10 billion tokens. Hmm. Yeah. Interesting. Yeah. You can count it. 10 billion tokens is about 10 billion watts. Yeah. It's, you can imagine that after the duplication, but you have at least 30 or 40 billion tokens to have finally 10 billion tokens in the domain corpora. I think it's not so easy.
1:01:32It's, it seems to be easy for many, many, but when we start counting them, it's not so easy to get 10 billion tokens. And I assume that that sort of what you're counting probably doesn't include like individual employees, email histories, all that sort of stuff. That stuff is like kind of out of scope. I think the emails, when they are emails between the, for example, the, the sales forces, for example, the, the, the, the, the call center or the sales emails, you can use them.
1:02:03But for example, there are also some kind of emails that are not keen to be used because of some undefined problems with the intellectual property or the risky cyber security risks and so on. I think it's not so easy to use any kind of email because always is the, the risk that, that some emails are too, I would say too, too risky or too, not too delicious to be sensitive. Yeah. Yeah. Sensitive. Yeah. That's really interesting. Yeah. That also does sort of make me wonder if new organization structures
1:02:36are going to be advantaged in some of these dimensions, because I totally understand the difficulty that would, would arise if you said, okay, hey, everybody. And we know you've been working here for all these years and sending all these emails, by the way, we're going to take all that and put it into our training process. You might have a revolt. It's a bit different that there is a, most of them, even the big organizations, they are not aware about what kind of data they have.
1:03:09And that is, before you adapt the AI or you be, before you train the AI, you have to, first of all, you have to clean your data stores, identify what data you have. What is there creepy or not creepy, the high quality or low quality. And that is the data curation process or data organization process. Everything that is, what is, what is, what is, what is around the topic, how to organize the data to
1:03:39find the, the high quality fraction of the data. I think it's the problem itself, yeah, that the companies, very often the company try to integrate or deploy AI or even train the AI without this data curation, data organization process. And it's usually, it go, it collapse. Yeah. I think this is the most important to be aware about what kind of data inventory you have, what kind of data level of qualities, what kind of data you can
1:04:10use without any regulations, internal regulation, external regulations. And this is the most important step. After this step, when you have the properly identified data sets, good, described and well organized, you can go up to invest in the AI training and AI deploying based on that. That's really all very interesting. Thank you. Definitely great food for thought for me. What about, just going back one more question on models. I know you had mentioned that some of the Chinese models have licenses that don't allow you to use them
1:04:45in the EU. If that… Mostly there were the LAMA 3.4, the first model in the license. There is a prohibition to use them in the, they are forbidden to use them in the European Union. But now there are Kimi models, the Chinese models, they have the same. They are not able to use them in the European Union. I think it's the problem with the AI Act. Because the AI Act that was released, the second chapter was
1:05:15released in August 2025. They demand from you to create for the general purpose models, the model, the cart of model. What data were used on training, how it was secured, what are the data sets used and what are the resources used for creating them. And so many, many other points in the model card. And I think those kinds of models, they don't want to validate what the data was used by them and what kind of even the training stages look like. Yeah.
1:05:49If that weren't a problem, would you be open to using Chinese models or are Chinese models not appealing for other reasons? I think it depends on the task you have. For example, I will be very, as you say, the sensitive, I will have a huge aversion to the risk when I would like to use the Chinese model to create the long structures of histories or essays and emails and so on. Because I think that when you create the longer forms of texts, the ability that the censorship's
1:06:24evidence will be, how to say, easily noticeable, is going up. But when you, for example, use some model for the task like the understanding, analytical extractive, for example, to extract key information from the documents. I mean, I am open to use the Chinese model because the risk of the censorship in the tasks that are typical analytical extractive ones is very low. Yeah, that makes sense. Turning to compute and talent, the other two big legs of the AI stool, I guess,
1:06:56how AGI-pilled would you say the Polish government is? I mean, it's pretty remarkable that all this is going on at the governmental level already. I would say that speaks to a pretty situationally aware and generally agile government. But how, you know, how committed is the government or how how big of a deal does the government understand this to be? And, you know, downstream of that are going to be questions around, like, how much funding is
1:07:26there? And is there? There are different, as you mentioned, during the first minutes of our speech, for example, there are three pillars of the AI revolution. The data, as mentioned before, the data organization, data curations, and generally looking at the data as the crucial point for the AI training. Next, the compute powers, the GPUs and the AI factories and others, the DC centers, the data centers and so on, and the talents, it means the people. Yeah, the three pillars, they together combined create the fuel for the
1:08:01AI revolution. And as I mentioned, I think in Poland, we are focused currently mostly on the AI factories. It means to buying as much as we can the GPUs and creating the DC centers for GPU tasks. Of course, we have some projects like the PLUM, that is a very good example, one of the, I think, maybe there are three or two projects similar to the PLUM in the European Union, that the Ministry of Digital Affairs, they are funds, they have funds, and they funded the consortium of the universities and institutes that are able
1:08:39to develop the models and the competences and so on. And somehow they support the talents in this grade because they have the money for the people somehow. But I think we don't be able to be competitive against the US market. Because in the US market, the AI engineers, they are paid like the NFL players. I hate something like that. The best AI engineers or AI researchers, they have the contracts like the quarterbacks in the NFL. The athletes are the stars. And I think we don't have such a maturity inside the decisions to
1:09:16pay people, to overpay people for their niche competences. Because I think it's much harder in the European Union, especially in Poland, to say what is the final objective function we would like to get. For example, 10 million customers, or I don't know, the $10 million each week in our subscriptions. It's much harder for the public sector to define such objectives that are very easily monetized and very easily evaluated. And I think in the USA, they pay so huge contracts because they are able to evaluate in some way,
1:09:51even making some margins, buffers for the future, how this talent can give you what kind of innovation and how this innovation will pay you back. And I think this is the problem, that in the USA, everything is there. As I mentioned, when I mentioned, when I was in Las Vegas, in ETLS, when I mentioned that we are working on the ALMs in the public ecosystem, that the Minister of Digital Affairs, they funded us. We created the consortium and this is funded by the public sector.
1:10:26Most of the people I met in Las Vegas, they were surprised that the public sector invested in ALMs. But in USA, it's hard to imagine that the public sector has enough intuition, enough knowledge, enough money to invest in such a sexy and very revolutionary topic like AI. So how do you think this will evolve over the next couple of years? Obviously, the amount of resources that the frontier companies are putting into their current and future models just continues to
1:10:58grow, right? That's expected to be. Even now we see that. We look at, for example, on the GPT models. And when you compare them, of course, the reasoning abilities are going up in different kind of models. But generally, the GPT-5, it was not such a huge, there's not a huge improvement over GPT-4. Of course, there are some reasoning abilities, but generally, there is a plateau. The models are going up, but since some level, they are going, the improvements are very steadily.
1:11:32There is like the horizontal improvement, not the vertical one. It isn't going up, but much more plateau, plateau, and their development is not so sexy as it used to be. Yeah. Remember that Chad GPT 2022, Chad GPT-4, the multimodality 2023-2024. There are some places in our history of the AI revolution that they were so shocking. And they also, they are so making our imagination work. Each year, we have something huge, something changing the rules of the game.
1:12:04Now, I think the models are going much more steadily. There are not huge bumps. But generally, I think that now we are starting to count the costs. The costs of the energy and the costs of what the models are used for. There is the place of the, I don't know how to say, the same thing is meeting people. That when you meet someone, or for example, you have some five minutes that you will be adored, or maybe
1:12:37not adored. And the same, after all, you have to evaluate, is it worth to meet with these people or not? Maybe some five minutes, maybe it's hours. But generally, I think now is coming the time where we are trying to evaluate what is the real cost of these tools. And how we can use them. And if we use them properly, what they give us. Yeah? Like the verification stage. Yeah? And I think the verification stage will give us a lot of information that
1:13:12we don't need to have very huge ALMs. We have to invest in the smart localized ALMs. Especially when you are working on previous solutions. Yeah, I sort of have mixed feelings about that. I mean, on the one hand, I do think already the models that exist are amazing artifacts for one thing. And very often, especially if you do take the time to do the supervised fine tuning and really dial in their performance, they can totally work perfectly well for all sorts of use cases.
1:13:47At the same time, of course, we've got the leading companies saying, you know, we're nowhere near done. This is definitely going to keep going. You should expect, you know, more progress. We're going to have AI scientists. We're going to have AI, you know, AI researchers. How much do you think your strategy depends on or will change if it turns out that there's not so much a plateau, but that you do still see, like, significant capabilities jumps, albeit with, like, you know, even exponentially
1:14:22more resources required to achieve them? Like, how do you think you navigate that world if it really is the case that a $10 billion training run, you know, is actually that much better than a $1 billion training run? I think the most important issue is, first of all, the demands and what is expected by our customers. Because I also mentioned that we are very often biased by the benchmarks, general purpose benchmarks, but they don't match
1:14:52the benchmarks and expectations that the business has. And, for example, if you have then your business, the company, and you know that, for example, you would like to use the AI, and in these places, we say, I should have at least this kind of metrics. It gives you the very good information what kind of benchmarks you should create to evaluate what kind of model is able to reach these benchmarks, these expected metrics. And I think this is the most important, that very often we analyze the general purpose benchmarks, which are, for
1:15:27example, the factual, the reasoning, you know, the extractive input competencies they are evaluating there. But, for example, for the business, the problem is slightly different. For example, they need something that write that beautiful email to the customer or something that makes the email that enables you to the cross-selling. That very often we don't know what should be done properly because there is no business, they don't define the requirements very explicitly. I will start from there, first and foremost.
1:15:57What would we like to improve in your business? What kind of tasks you would like to send to the AI? Next, create the benchmarks for these tasks. And next, you choose the airlines. Because I think most of the business cases I have seen, they don't demand reasoning. You can do it with the normal airlines without the reasoning stages. I think we should try to, we should slow down a little bit and analyze what is needed to be done. And what is especially because the huge business factor, not only because it's a sexy and public relation likes
1:16:34that, but what gives the money to the business? Or what makes some savings? Yeah, there's a massive disconnect, I think, often between general business culture and the culture of things. There are two trains, but they are not in the sequence, one after one, but they are next to each other. One is much faster, the second is going on their base, but they are not, as I mentioned, very often there is no crossing between them. There are two roads, but the crossing is very far before us.
1:17:07Yeah. Yeah. Yeah, that's really interesting. Who are your allies in this? You mentioned, you know, using multiple languages, and I assume that that's sort of in some sort of partnership with maybe other neighboring country national institutes. I guess I'm curious as to how you think the international dynamics will play out. Historically, in the Cold War, we had, you know, the US and the USSR, and, you know, these two great powers were sort of engaged in proxy conflict and whatever all over the world.
1:17:39And a lot of other countries understandably said, you know, this is bullshit from our perspective. And there was a movement of countries that were like, we don't really want to be in either of your camps. We would rather be just independent. And, you know, the beef that you guys have between yourselves, like, we don't really want to be, you know, a pawn in that game. Now it's the US and China, obviously, that are kind of the two, you know, big poles of AI power.
1:18:11How do you think countries that are, I sometimes say countries three through 193 on the AI power rankings, how do you think they will react? Like, well, do you see alliances forming or, you know, countries working together to share resources, share data sets to try to create some sort of a third way in the AI space? There are some movements in the European Union, for example, there are projects that are international. They gather the different kinds of people from different countries and try to do something together.
1:18:44But I think it's a huge problem that generally when you would like to get very fast products or very fast outcomes, you have to centralize. The problem is always the same. Generally, the best way is to have the federation. Everything should be spreaded out. You have different people in different countries. They collaborate each other. They build the, how do you say, the wealthiness is going up everywhere. In somehow distributed way. But in a normalized way. But generally, when you would like to get very fast incomes,
1:19:16very fast outcomes, and have the products in months, not in years, usually you have to centralize the assets in one place. And there's a problem. Because there are two different ways. If you would like to do it in an ideal way, you should create the unions. The unions of countries, the unions of states, the unions of the partners, the networking, the consortium with the consortium with the plenty of hundreds of stakeholders to get this knowledge everywhere and to distribute this knowledge and this power
1:19:48everywhere. But usually, if you have to get some outputs very, very fast, you have to centralize. The problems are, there are two opposite ways. You're not able to do it the same way both of them. And I think this is the problem. Yeah. Because generally when the, for example, when you have some very huge pressure on the outcomes on new models, you always prefer the centralizations. Yeah. Like with the Silicon Valley. Yeah. You have the huge USA, but most 90% of the startups in Silicon
1:20:20Valley. Yeah. The centralization is one place where there are the money, assets, and everything. But from the economic point of view, the best place will be to distribute these companies across the whole USA. Yeah. And I think that is the problem. Yeah. That if you are, if you would like to monetize something and you would like to get the very fast outputs, you have to centralize. But generally the best for the economy and social aspects is to distribute in a normalized way across the country
1:20:51and across the continent. What about geopolitical? I think there are still two players, the Chinese and the USA. They have the two biggest economies. They have money. The Chinese companies, I heard that they pay the same money to the researchers as the USA ones. Yeah. The contracts are now currently, they are somehow similar. It means they pay very well. They don't have to compete. There is no risk that they will be taken over by the USA companies because they are well paid in
1:21:22China. In Europe, I don't think there is a still there. There is a Mistral, the European funded startup. Now it's not the startup, but it was the startup two years ago or three years ago. But now I heard that the shares in the 30 or 40% of the shares are in the Microsoft. Yeah. There is not so open and as it used to be because there are some stakeholders from the USA. I think generally the problem is slightly different. The question is whether the Chinese and USA, they are going to be the rivals still, or there also there's
1:21:59a chance for the cooperation. This is the question. Maybe it's a chance for cooperation still. Yeah. From your lips to God's ears. Maybe just one little follow up, and I think this has been excellent. I really appreciate all your time and all these thoughtful answers. Is there anything that you have seen that on sort of a technical or socio-technical level can help with the cooperation of decentralized AI? Here I'm thinking about things like the NEAR protocol. I recently did an episode with the creator of the NEAR protocol, Ilya Polosuhin, or there's also the intelligent internet,
1:22:34which is Imad Moustak's project. There's others as well. These things sort of have this idea that if we create the right scheme, it might be somewhat cryptographically enabled. There is a topic called federational learning, yeah, federated learning, that you can use the different data sets, somehow anonymize and secure, and use them in a way that in networks that you are able to identify the necessity of data, but you can use the data to train your models. Yeah. There are different kinds of ideas about this federated learning.
1:23:05But generally, I don't know that there is some huge deployments of such approaches. Yeah. Even though there is a very good approach, yeah, to have such networks or federations, to cooperate with each other and to share data in some, as I mentioned, the secured way. But I think still there is, we are not on, we, if we talk about the business and economy, we are not on the level to use it. Yeah. Because I think we are still on the problems that we, the companies mostly, they don't know what data they have.
1:23:39What is the quality of their data? What is the value of their data? If you aren't able to, to measure your own in-house data repositories, how you can go further and create some data mixture or the, or the network of data, data repositories. Yeah. I think this is the, maybe in the future. Yeah. In the future, I think there is a chance that there will be the, the, like the, the distributed data repositories through the secure levels and anonymizations used by the huge consortia.
1:24:09Use the, as much data as we can. But I think this is not the, the, this is not for the next year. Yeah. Yeah. So much depends on whether there really is a plateau or whether the frontier companies are just going to continue to keep scaling successfully. Yeah. And now I heard that this, the plateau is caused by, because there is a lack of the organic data. Yeah. The biggest companies, they, they collected already all, almost all organic data that was able to be collected from
1:24:40the internet. They, even some of the companies, they, they, they, they scan the books that they're not, that they haven't appeared in the internet to make this organic data gains relevant. But there is still the problem that I think that there is the, this repository, the resource of the organic data is almost full. Yeah. That's why we're now seeing all these simulated worlds and yeah, the strategies to overcome that are, are going to be definitely fascinating to watch. I for one will bet on them working, but it does to a certain degree, certainly remain to be seen.
1:25:15Um, again, really fascinating conversation. It's been awesome to get your perspective. Um, anything else you want to share or anything we didn't touch on? You want to, you want to comment on before we break for today? I can share that too. I can recommend to our archive paper, The Plume Family Models. It will be the archive, at least now the archive archive paper released in November, the first days of November. And I recommend readers and the viewers of this podcast to look inside this paper.
1:25:45The Plume Family is the archive, the title of this, of this paper. Yeah. That's P-L-L-U-M. P-L-L-U-M. P-L-L-U-M. Yeah. And this, uh, the, the archive paper, almost 100 pages. And the P of course is for Polish. It's the P-L-L-U-M. P-L-L-U-M. Merrick Kozlowski. This has been amazing. Thank you for being part of the cognitive revolution. Yeah. Thank you very much. If you're finding value in the show, we'd appreciate it. If you'd take a moment to share it with friends, post online, write a review on Apple podcasts or Spotify,
1:26:17or just leave us a comment on YouTube. Of course, we always welcome your feedback, guests and topic suggestions and sponsorship inquiries either via our website, cognitive revolution.ai or by DMing me on your favorite social network. The cognitive revolution is part of the Turpentine network, a network of podcasts, which is now part of a 16 Z where experts talk technology, business economics, geopolitics, culture, and more. We're produced by AI podcasting. If you're looking for podcast production help for everything from the moment you stop recording
1:26:47to the moment your audience starts listening, check them out and see my endorsement at AI podcast dot ING. And thank you to everyone who listens for being part of the cognitive revolution.
More from The Cognitive Revolution

AI:AM Highlights: Zvi on Pacing & Trump-Xi, Astra better behaved than Fable? + a new LLM Pain Axis??
Sep 19, 20261h 41m

No Code Is Code: Zapier CEO Wade Foster on Headless Tools, Zapier MCP & Automation Bench
Sep 17, 20261h 8m

The Balance of AI Power: Anton Leicht on Politics, Pacing Deals, and Muddling Through Well
Sep 15, 20262h 10m

AI:AM Highlights: Astra as AGI, OpenAI's Pause, Mythos @ Mozilla & Human Agency vs Technocapitalism
Sep 12, 20261h 42m

Nathan Goes to China #3: US-China Relations, the Art of the AI Deal & the Road to Pax Robotica
Sep 10, 20263h 17m