The Festival of Genomics and Biodata is just around the corner, and we’ve had the opportunity to sit down with some of our expert speakers to get a sneak peek into what they’ll be discussing, and why they think you should come along to the event.
In today’s interview, we speak to Alvis Brazma (EMBL Scientist Emeritus, EMBL-EBI) about the importance of data sharing, the complexities that arise when handling sensitive data and the need for collaboration and standard practices.
Register for the Festival of Genomics and Biodata here.
Please note transcript has been edited for brevity and clarity.
FLG: Hi everybody. Today, we’re joined by Alvis Brazma from the European Molecular Biology Laboratory. Alvis will be joining us in January at The Festival of Genomics and Biodata for a really exciting talk, and we’ve got some time today to chat to you and find out a little bit more in advance. So Alvis, over to you. Can you tell our audience a little bit about yourself and your career?
Alvis: I was born in Latvia, which at that time was still the part of the Soviet Union. I studied mathematics there and I got my PhD there in computer science. When the Berlin Wall came down, I took the opportunity – I spent almost two years at New Mexico State University, and then another two years at the University of Helsinki, part time, and I joined the European Bioinformatics Institute of EMBL in 1997. There, I became a member of senior faculty in 2004 and I was involved in establishing quite a few core resources, some of which are still run at EBI, and some of which have been wound into others, including ArrayExpress, which was the first database for sharing microarray data. Then Expression Atlas, which was an added value database to build gene expression patterns on top of what we were collecting in ArrayExpress. Then BioStudies database, which took over from ArrayExpress as the main database for functional genomics data at EBI. I was also involved in establishing BioImage Archive, and in parallel, I was running my own research group.
I’ve been involved in quite a few international consortium dealing with big data. I think probably the best known is the International Cancer Genome Consortium, ICGC, where I led bioinformatics in some of the projects.
FLG: You’ve been involved in a lot of really exciting projects, then. It sounds like you’ve had a really varied career, and you’ve been to a lot of places. It sounds great.
Now, obviously you’ve been involved in this kind of work for a while now. How have you seen the landscape of data sharing evolve in the last couple of decades, and particularly when we’re talking about sharing patient derived data?
Alvis: The change has been enormous. It all started before patient data was a major source of data. And then in the early 2000s and late 1990s the default for what most people, most scientists, would think, was, ‘well, no, I’m not sharing my data.’ In the last two decades, it has changed completely. And I guess the big push was the Human Genome Project and the Bermuda Agreement that the data should be released, but that, of course, was applied only to human genome data. So, when the microarray data came out, we had to start it all over again. Few wanted to share their data, and even if they would share the data, it was not in a way that was usable for others.
So, we started this microarray gene expression data society and established this minimum information about microarray data, which is known as the MIAME Standard. And I think that was the very first standard that explicitly said that the data on which you base your scientific research, on which you base your publication, have to be made available for scrutiny, and also for others to build their research on top of it. That was accepted, and that transformed, I think one could say. Although there were other sources for ideas [leading to] that is now known as FAIR principles (findable, accessible, intolerable, reusable), but effectively, MIAME established these [principle] for microarrays at the time [first], and now it’s broadened. But that was black and white. And I must say, when it comes to these data, let’s say from model organisms, I’m pretty much a data sharing “absolutist”. If you publish, you have to share your data. Of course, you don’t have to publish, and you can do whatever like, but if you want to publish, you have to share the data.
When it comes to patient data, this complicates things enormously, because there’s a good point to be made that the data derived from me, collected about me, belongs to me. I think it’s a very strong point. I would need to consent to this data [being shared]. Maybe I want to consent only for some specific research. Certainly, I don’t want to consent to it being abused. I don’t want to consent in a way that can later come back and, for instance, I’m denied health insurance, so my premium goes up. Things become much more complicated immediately. Again, I think when it all started, there was this fact that it’s tricky, and that there’s no black and white anymore, and that was used by many not to share their data at all.
And again, there were some initial projects – I think Wellcome Trust had the genotyping project and they established what was called the European Genome Archive, now it’s named the European Genome Phenome Archive (EGA) [at EBI] – to at least to give means for people to share the data. That’s different from public data, because to actually submit data there, you have to have consent from these patients or human research subjects, and the consent may be specific on what data can be used. Then, you need a data access committee, so if scientists want to have access to this data, they have to submit the proposal to the data access committee to evaluate if it complies with the consent.
So, that’s how it all started, and it’s still a bit like pulling teeth. It’s hindered by many factors, but attitudes are changing. I would say it’s now more and more accepted that, even if it’s human research subject data, where you have all these privacy concerns, mechanisms need to be found as to how this data too can be shared and made usable by others.
FLG: Yeah, it definitely sounds like it’s quite a complex issue. It sounds like, on the one hand, it’s really important that the data is shared, but then there are some very genuine concerns that people have, and I imagine that really adds to the complexity. I do find that interesting – and it all makes sense when you discuss it from the perspective of patient-derived data – but what were some of the reasons prior to the Human Genome Project that people were apprehensive to share that data? Were there any scientific or technical reasons alongside that?
Alvis: Well, I think there were a lot of different reasons. Perhaps one of the most [important], I must say, from my experience, was that people invested a lot of energy and effort to collect the data,. You know, when epidemiology was kind of the main source of data, epidemiologists invested lot of effort, a lot of their careers, into collecting this data. And they thought – and to some extent they may have been right – that if they make this data available to everyone, then they can lose that asset. And that was a bit different from, let’s say, high throughput data, where it’s actually not that difficult to generate the data, really. So, we can understand all these reservations. But the thing is, then, why are you doing research? Why are you collecting this data? I think, generally, everybody does it for the benefit of the human good, really. Of course, you also need to earn your salary. There’s no way around that. But still, if you don’t, in any way, release this data, then what’s the point? It was often that you would find a collaborator who gives you something in return, like putting your name on a paper, for instance, and that was certainly the case then, and to some extent it’s the case still. People find all kinds of reasons not to share their data, unless they get something back for it, which is to some extent understandable. But on the other hand, most of this research is publicly funded. So, the public funding agencies – well, many of them, like the Wellcome Trust, for instance, or NIH in the US – are pretty strict. All the data that is generated by public funding should be made available in a reasonable way.
FLG: Yeah, absolutely. I think you’ve highlighted there how important is that we share data, but it’s also really important that we share it, specifically, internationally as well. What is currently hindering that?
Alvis: Well, it’s fragmentation, really. Different countries have different legislatures. Different countries and even different institutions have established different infrastructures that are not interoperable. There’s a lack of standards. So, everything that is difficult within one country becomes exponentially more difficult with every new country being added to a consortium. But it’s very important that it happens. A very specific example that shows how important it was that data was shared internationally was COVID-19. Only because scientists did agree to share this data internationally, we were able to figure out what actually was going on, how the disease spreads, how it evolves, and also, eventually, to develop an effective vaccine, and an effective vaccine for every new variant that appears. Without international sharing, that’s totally unimaginable, really.
FLG: That’s a really interesting point. A lot of people are highlighting that, when they discuss things like this, the COVID-19 pandemic really changed a lot of people’s views. Are you seeing that spill over into different fields, that sort of collaboration, and reciprocity, I suppose?
Alvis: Well, it’s a very slow process. It’s interesting, the COVID-19 pandemic has been often treated as something very special. And I noticed that a lot of this is already forgotten. I think it gave a push and people better understood why [data sharing] is important. But that it had a major spillover effect, let’s say, on cancer research? I’m not sure. Although, I haven’t been involved in hands on cancer research for the last five years, so maybe I’m wrong, but my impression is that there was a push, it helped a bit, but they are now back to the “normal” flow of evolution, as it was before.
FLG: That’s really interesting to hear, and it sort of brings me on to the next thing I was going to ask, which is about cancer genomics projects. Now, I imagine there are probably some pretty unique challenges associated with those projects. What are those, and how are they being addressed at the moment?
Alvis: On one hand, sharing cancer data, in some sense, is easier than many other types of data, partly because the patients – probably due to the high lethality of this disease – are quite willing. When they donate their samples, I’ve never seen anybody who has said, ‘No, I don’t want this to be used for research.’ But one of the challenges is, again, what is the consent? Is there consent to do research on the data for that specific cancer only, or any cancer? I remember during ICGC, that created some complications.
I think what is specific to cancer, and that kind of spills over to data sharing, is that these data are very complex. When we talk about rare diseases, let’s say rare metabolic diseases, it’s one mutation in a certain enzyme and all you need is to find that mutation and you know what’s going on. In cancer, it is a genome disease. You don’t have one genome anymore. You have your germline genome, which you inherited, but then you have your evolving cancer genomes. We typically try to simplify this by having your germline genome and somatic genome, but that somatic genome is just a snapshot of something. You have this evolutionary trajectory there. And then there’s also heterogeneity within one patient. You need the images from pathology, it all needs to be linked to the data. Now, with single-cell data, you have many, many different cells, and there is spatial omics. So, it’s the complexity, really. Cancer has turned out to be a more complex disease than was hoped, let’s say, 10 years ago, when genomics started having an impact there.
FLG: It definitely sounds like it’s a challenging area to work in. That heterogeneity really must hinder things, I can imagine. So, what are some of the key considerations to make when building some kind of federated data sharing infrastructure, especially concerning the security and the privacy of this sensitive data?
Alvis: Federation has sometimes been given as the answer to privacy, to some extent. I think that the consideration is that first you have to ask, why? Federating is actually more difficult than centralizing. Perhaps not necessarily in terms of funding, because the funding agencies don’t want to give a lot of money to one place. So, in terms of funding, maybe actually it’s easier to fund federated infrastructure, particularly internationally. But in terms of technical development, it doesn’t simplify anything, really.
The project I was involved in when I was heading the Omics section at EBI was the Federated European Genome-Phenome Archive. I haven’t been responsible for that project, I haven’t done very much there, but I was kind of like a mentor for people who were running that project, so I know quite a lot about it. [This project] was conceived to solve the problem that some of the countries had legislatures that didn’t allow patient data to leave the borders of that country. Some countries have more fuzzy legislature, and then sometimes researchers simply use this fuzziness to say, ‘no, no, the data can’t leave our country.’ So, the people within EBI and our collaborators in CRG Barcelona, we decided to solve that by Federation. That kind of solves that particular problem, but it doesn’t solve all the problems, really. So, suppose there is a cancer data set in, let’s say, Germany, in federated EGA, and you want to do research on it. How do you do that, if you can’t download the data? You need a secure environment, cloud infrastructure, and you do your research there. But then you still have to take something out of it. What is it that you can take, and what you are not allowed to? And then you can’t combine that data set with another data set, which you have in your own country, for instance. So, it does solve some problems, but it doesn’t solve some other problems, really.
The main thing, when I started thinking about this question, is, how will it work in the new age of artificial intelligence? Because all this infrastructure has been set up for people finding the data sets, they’re doing the analysis. But now you want to do it, or soon you will want to do it, much higher throughput. You will want algorithms to do that, rather than people. How Federation will adjust to that, I don’t have an answer. I’m sure people are thinking about it, but I can’t say that I know how it can be done.
FLG: It seems like AI has really changed the landscape when it comes to addressing these questions. And it seems to, at least from the perspective of somebody like myself, who isn’t in the weeds doing this work, it seems like it’s coming really quickly, so that brings some challenges with it.
Alvis: Yeah, absolutely. I know that at EBI, we are now thinking about how to make the public data that we have AI-ready, and a lot of effort is going into this. I think we will solve that problem. But when it comes to this controlled access data, which is very often federated, it’s two orders of magnitude more difficult and a more complex problem. We are not, I think, close to having a solution.
FLG: And I suppose even if AI is the solution to all of these problems, it throws into question all the work that came before it, which must make things very challenging.
Alvis: Of course, a lot of people are also worried about abuse of AI, like any abuse of this private data, and you cannot ignore these problems. They are very, very important. You need appropriate legal frameworks for that, you need established traditions of what can and cannot be done. But unless we do open this data for AI use, the full potential will not be used.
FLG: Yeah, yeah, absolutely. It’s definitely a bit of a minefield tackling AI at the moment.
How do you see the role of international collaboration and building and maintaining these federated infrastructures, and are there any current initiatives that you find particularly promising?
Alvis: There is an ongoing initiative that has done a lot, and still I think is the main one, the Global Alliance for Genomics and Health (GA4GH). I think they have hit on the head of the nail, really, by focusing on developing common standards and common understanding. It’s impossible to build the world’s database on this and that. In some sense GenBank or the European Nucleotide Archive is a bit like that, but that’s dealing with much simpler and much less data, in a way. For patient data, you won’t have one database where you have everything, this will have to be federated. But that effectively means that parts of the federation have to be interoperable. And the role of these international consortiums, one of the most important roles, is to build these standards that allow interoperability.
But then there is, perhaps, a softer thing there, which also should not be ignored. This federation, in my experience when it has worked, it has worked because scientists in different countries, or in even the same country, had come together and established trust. This all starts working when the group of scientists working on the infrastructure trust each other. I think that all of these international consortia are bringing people together so that they can talk to each other, and they build this trust in each other, and also build a common language and understanding, which then can translate into more formal standards. It allows this bottom-up building.
So, I think GA4GH is what I would mention first. On a smaller, more focused, but perhaps more practical, scale is Federated EGA. In between, maybe, there is this European Union-funded project, Beyond One Million Genomes [(B1MG)], which also is making a lot of advances, and I think will be very important in building these things.
FLG: It sounds like there’s a lot of work going on to tackle these issues and advance research, which is really exciting.
Alvis: Absolutely.
FLG: You’ve already somewhat touched on this. But how can researchers navigate the challenges that come with this kind of work, while maintaining data integrity and especially patient privacy?
Alvis: On one hand, it’s important to keep in mind what the worst-case scenario would be if something goes wrong, and be aware of it. But that should not make you too cautious, because if you look at the reality, so far, I’m not aware of a single case where there has been intentional [genomics] data abuse. There may have been some honest mistakes, releasing data that was not supposed to be released, but even then I’m not aware of a single case where somebody would have suffered from that, really. This kind of over-cautiousness is bad.
At the same time, you have to be aware of it in a way. You have to think about the worst-case scenario. For scientists, it’s very difficult to really judge that. So, that’s where the law should come in. And if there’s a law, you need lawyers, and lawyers very often are over-cautious. They are kind of trying to close all the doors so that nothing bad can happen, and that can hinder the research. But I must say, in my experience, working particularly in cancer but also in some other consortia, I have come across very, very good lawyers who understand that there needs to be this balance, to allow scientific advances as fast as possible, whilst at the same time protecting the patients. And what I would say is that scientists should not be afraid of this, and they should build consortia that are big enough that you can afford to involve these lawyers, because it cannot stay as a cottage industry anymore.
FLG: It’s interesting that there are all of these layers. There’s the scientific side, there’s the legal side, and it’s all just as important as each other, and there’s so many different people in one room trying to build this infrastructure.
Alvis: Absolutely, yeah.
FLG: And as an advocate for data sharing, what advice would you give to institutions or individuals hesitant to share their data due to some of those concerns or potential misuse?
Alvis: I would have to say the answer is the same as to the previous question. Maybe I strided into this question already! Really, I think that the default is we want to share the data. If your people tell you, for example, a head of the institute, ‘no, we can’t share the data,’ try to understand why. Is it a real, genuine concern, or is it just that it’s more convenient not to share data? Maybe we can get more credit, or something. It almost never works like that. So, if there are some genuine concerns, then I think you need to see the legal landscape, what you can and what you genuinely can’t do, but your default should be that we want to share as much data as possible, as easily as possible.
FLG: Yeah, that’s good advice. And how can we balance the need for open data sharing with that need to protect sensitive information – particularly in the context of cancer research, where it’s so complex?
Alvis: That’s probably the most difficult question of them all when it comes to data sharing, and I don’t think there’s a straightforward, simple answer. You have to, in your particular case, with your particular data sets, with your particular collaborators, have a very honest discussion about what harm can occur there, what is over-cautiousness, and how can we make the data as widely useful as possible. Probably, each case is a bit different, I’m not able to give one simple answer here. It’s just the same thing, that your default should be ‘we want to make this data as useful as possible for everybody in the world,’ really. But then does our legal framework really allow us to do this? You need good lawyers, probably, as well. Good ones, not the one who, by all means, wants to protect your institution, but somebody who understands why we are doing all this.
FLG: It sounds like a lot of the issues are not really scientific or technical. It’s about people. It’s social and ethical, and ultimately, most people are probably trying to get to the same place in terms of helping people.
Alvis: I think the technical developments are part of the solution, but they’re not the most complicated thing in my experience. Then, of course, there’s very fancy cryptography-based [technology], where you are allowed to search data, but you can’t ever see the data. I haven’t seen them work, but I also don’t see what problems they are really solving. The technical solutions really are about how you encrypt your data so that it cannot leak, and these are not super difficult. Then federation, how do you build a secure environment, I don’t think that’s where the difficulties really lie. The difficulties lie in the social aspect of all this, and that inevitably leads also to legal questions.
FLG: Yeah, it’s really interesting to look at it from that angle. What do you think the data sharing landscape is going to look like in the coming years?
Alvis: Well, I think there will be a slow evolution, and there will be a lot of pressure from people who want to use machine learning and artificial intelligence to get access to this data, which perhaps will accelerate the progress. I don’t know whether it’s a good or bad thing, but a lot of funding, including private funding, goes into this, and that will create all kinds of pressures. Maybe there will be some breakthroughs, it’s very hard to say. But when it comes to international data sharing, it has moved very slowly, and for a good reason. I can’t see that suddenly there will be one thing that will change it a lot, really.
FLG: That’s interesting to hear. Moving away from that, can I ask you about what you do outside of your work at EMBL? Because I hear that you have released a book relatively recently called Living Computers. Could you tell us a bit about that?
Alvis: Thanks for asking! It was partly outside the EMBL, but a part of what I was doing there. The full title of the book is ‘Living Computers: Replicators, Information Processing and Evolution of Life.’ It was published by Oxford University Press about half a year ago.
Why did I write it? As I said at the beginning, I was trained as a mathematician. I come from that mathematics, theoretical computer science background, but then for three decades, I’ve been working with biologists, mostly at EMBL. I have been incredibly lucky to be in the scientific discipline called bioinformatics, which combines both the mathematical side and biology, almost since the beginning of the discipline, when I joined EMBL and became a bioinformatician, it was just a fringe science. Now it’s everywhere. One observation that I had during these three decades, looking at how things develop, is that the concept of information was becoming increasingly central to life sciences. So, having all this experience and having made this observation, I decided to write a book explaining biology and the phenomenon of life from the point of view of information processing. I would say, if you try to understand why information – and information processing – has become so central to the life sciences, you come to a conclusion that, effectively, life and information are two sides of the same phenomena. There is no life without information, without genomes. The genome is information. Arguably, the very first recording of information and processing of information started with life. And, perhaps this ability to process information really is one of the defining features of life.
But then came computers over the last 50 years or so. For the first time, we have the ability to process information outside what we normally call living systems. Suddenly, we can do it in silico or using all kinds of other media. To me, it was an interesting observation. And I thought, now we cannot just process information outside the human brain, the living cell, we can also learn. Computers can learn now, with large language models and deep learning. I think, to me, the question that arises is, will this lead to completely new life forms, which will evolve in their own way? So, I don’t have answers to these questions, but for those who are intrigued by them, I can recommend you to read my book!
FLG: It definitely sounds really exciting, it’s an interesting question to consider. I’m a bit of a geek when it comes to this sort of thing, so I think it’s really cool that we can now do things in silico and use computers in this way. So, I’ll certainly be giving your book a read, it sounds great!
Now, as I mentioned earlier, you’re speaking in January at The Festival of Genomics and Biodata, and we’re really excited to have you there. But what excites you about The Festival, and why did you agree to speak?
Alvis: I think it’s one of the main, and the most exciting, scientific events of the year. Certainly in life sciences. It brings together so many people with different backgrounds and from different companies, and with different ideas. I want to be there simply to look for new ideas, for inspiration and to broaden my horizons. I’ve always wanted to come to the event, and despite Cambridge and London not being very far apart, something always came between me and the event at that time of the year, so I never managed. So, when I got this invitation, it made me super, super excited, and my only worry was something may already be in the calendar! But it wasn’t, so I will be there, and I’m really looking forward to it. It will be very exciting.
FLG: Well, we’re looking forward to having you there. We’re also really glad that there wasn’t anything in the calendar! It’s lovely to hear you say that. And as I say, we’re really looking forward to seeing you in January. Thank you again for joining us today. It’s been really lovely to chat to you, and we’ll hear more from you at The Festival of Genomics and Biodata.
Alvis: Thank you very much, Lyndsey, it was pleasure to talk to you, and I’m also looking forward to the event. Thank you.
Register for the Festival of Genomics and Biodata here.




