The Bioinformatics CRO Podcast

Episode 94 with Michael Fanous

Michael Fanous, founder and CEO of FanousPhotonics, discusses imaging systems for pathology and Scanimus, a research-use only hybrid microscope scanner system.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Michael Fanous

Michael Fanous is the founder and CEO of FanousPhotonics, which makes AI-powered digital pathology scanners. Scanimus is a research-use only portable digital microscope and slide scanner for pathology.

Transcript of Episode 94: Michael Fanous

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. I’m Grant Belgard, and today I’m joined by Dr. Michael Fanous, founder and CEO of Fanous Photonics, which is developing Scanimus, a research use only imaging system. Michael’s research at the University of Illinois Urbana-Champaign and UCLA brings together microscopy and machine learning, including work on continuous scanning and automated tissue analysis. Today, we’ll explore what we should demand from a scientific image, how an idea finds its way out of the lab, and what he’s learned about choosing problems and building a career. Michael, welcome to the podcast.

Michael Fanous: Thanks a lot, Grant. It’s great to be here

Grant Belgard: How would you describe the work you’re doing now, and how do your different roles fit together?

Michael Fanous: Yeah, so I’m definitely juggling a lot of different tasks, different roles scientific, entrepreneurial. so on a daily basis, I’m transitioning from a design, scientific, engineering mentality to a more like financial, entrepreneurial pitch, sales, that sort of thing. becoming increasingly easy and smooth, but it’s still a work that I have to navigate through. And so it’s not always as tidy as I’d like, but I think I’m increasingly getting the hang of it and making things more efficient, understanding which efforts are helpful and, productive and which are just draining and time-consuming. So I’ve been getting better at making it tidy and, know, efficient.

Grant Belgard: What is Scanimus and who would you most like to see using it?

Michael Fanous: What is Scanimus? it’s a microscope, it’s a scanner, it’s a hybrid microscope scanner. It’s a category-defining product. I don’t know exactly what it may be called in the future if it really takes off. Maybe it’ll be called a scanoscope, maybe it’ll be called a micro scanner. We’re really trying to create something in between the two systems that are available to pathologists right now, which is the manual microscope, the centuries-old technology, and the modern scanner which is usually operated by someone else in a different facility, an involved piece of equipment. And so we’re really trying to bridge that gap. We find that there’s a great gulf in between the two worlds of the manual old tech and the modern scanner.

Michael Fanous: And we think there’s a real opportunity there to try and emulate what a pathologist is doing when they’re interacting with a microscope and try to digitize that process and facilitate and streamline that workflow. what we’re trying to do. And of course, the pathologist is our main primary potential customer. So that’s what we’re really trying to cater to is the pathologist first and foremost.

Grant Belgard: Could you walk us through one intended use case from a specimen arriving to a pathologist deciding what to do with the result?

Michael Fanous: So this is a research use only device. I should specify that. And really in terms of intended use, because the statistics in pathology are so modest with respect to how much is being digitized in the state-of-the-art top-tier clinics, institutions, what have you, are just trying to improve those numbers. So whatever the use case may be, assistive review, triage, finding a region of interest quickly, we’re able to intelligently navigate through using different objectives and magnifications and depth of fields the way that a pathologist might. so in just improving the numbers and the statistics and having a scan, irrespective of how it’s gonna be used, if it’s gonna be used for just save it later for further examination, maybe use it for one of the software companies, the training data sets. You hear about like millions of slides that they’ll have, and it’s considered like a boast.

Michael Fanous: But millions of slides is really scratching the surface. The human body has so much tissue. If you were to actually section everything, you’d have billions of slides. For one human an adult with sixty liters, seventy liters of volume you have way more information, and there’s a great deal to be done. I think we’re just really at the absolute beginning. This is a very embryonic. The whole digital pathology field is we’re at the inception of it. This is the dawn, I think it’s an exciting time, but it’s also a time to, I think, take action and be more proactive about digitization, making scans, and, more images and creating larger data sets. I think that’s in itself very useful.

Grant Belgard: Where do you see the most consequential bottleneck in turning a glass slide into useful biological information?

Michael Fanous: Yeah. The, the greatest bottleneck currently is just to take images. It sounds trivial and obvious, but if you’re not taking pictures, you can have all the AI algorithm and sophistication and power that you like. But if nothing is recorded, even at lower magnifications, even at cruder resolutions, then you cannot process that information. So having affordable, accessible, ergonomic, this is another important one, systems, portable this is a, a real big barrier right now that we’re trying to overcome with a system that is just very easy, very intuitive, and really affordable. This is undercutting a lot of the, the competition by orders of magnitude.

Grant Belgard: How does continuous scanning work and what trade-offs does it entail?

Michael Fanous: Yeah, so trade-offs are really important. So is nothing free. There is no free lunch. We pay for everything that we initially are able to compromise. We pay very heavily in training. It’s all about training. If you see what happens behind the scenes, you’ll realize that our system, yes, it goes very fast. We’re able to reconstruct rapidly, but our training is very slow. So let’s go through what is continuous scanning exactly in our case and how it compares with most scanners. Most scanners will use what’s called stop and stare. So the specimen is literally physically halted during an acquisition. There is no movement or else you would get what’s known as motion blur.

Michael Fanous: There are some special types of cameras that use a line scanner configuration where the sample is moving continuously, but you’re not seeing that interactively, and then there’s a post-processing event that computationally renders the image. And so that’s expensive. It needs to be synced with the stage. It’s involved. We are trying to circumvent all of that with standard, more modest microscopy hardware and use just a standard global shutter camera, CMOS, and have it interact, being displayed as you’re taking the scan. So the pathologist is seeing this. They’re looking at the slide move by, and they’re able to see the scanning process real-time, which already gives them an idea of what they wanna look at and what they wanna see. And so this is really helpful.

Michael Fanous: And then we use what’s called an image-to-image translation network to convert that motion blur image into its sharp counterpart, which sounds a little bit bold and maybe a little reckless. But because our data sets are carefully put together and all this is supervised, so there needs to be labels. So can’t just throw in data and hope for the best and make sure the algorithm just works out. No. This is very carefully selected in terms of hyperparameters, the trainings all of the data for each different staining type needs to be carefully picked out and hand-selected. And then, different objectives require different training types. So when we have our ground truth, we get that through slow scans. We can’t do that just using a normal stop-and-stare stitch. So it’s actually very time-consuming. It’s boring on our end for that part.

Michael Fanous: We’re just trying to make sure that the fidelity is as high as possible. So that’s what we’re doing, if that makes sense.

Grant Belgard: How do GANscan, BlurryScope, and Scanimus relate to one another, and what are the important differences?

Michael Fanous: So those are three different names for the three real big milestones of the evolution of this technology. GANscan is really proof of principle. It’s a theoretical, mathematical-based concept of converting the motion blur into the sharp image. We used at that point lab benchtop microscope. We fiddled in the background software, and were able to manipulate the stage to some degree. We were able to make it go fast, make it go slow, and then make an image to image. But we didn’t build anything. There was no concrete hardware assembly on our end. And so we didn’t pursue anything commercially or patent anything at that point.

Michael Fanous: When I then went to UCLA as a postdoc and we wanted to take it to another level, make it a little bit more dexterous in terms of how much we can manipulate the speed, how we can play around with a video that is not sync with the stage but increase the post-processing ability. We actually built something, and that was designed for immediate classification on images, comparing that, determination with a real pathologist assessment. And so we wanted to see how far we can take that. And at that point, there was a device. It was in the lab, it was sitting there, and it was, asking itself, “Can I be presented to a pathologist in a commercial setting?” To me. So that’s when I decided, this is really potentially something, and I had this notion that, pathologists in the West would find it very appealing if I had just added a few components. At that time, it wasn’t an integrated device.

Michael Fanous: You needed separate laptop, a separate computing device, And so I thought if it was a whole integrated system, which pathologists don’t have really on their desks, ’cause I had worked with many pathologists in my PhD program, at the postdoc. I had seen their offices, and I figured they could really use a nice, elegant, easy-to-use gadget where it’s all there, and they just need to push a button. It’s extremely simple to use. And that was the designing. The engineering was the same as BlurryScope really. There was no new groundbreaking engineering or science. It was just about design. And from BlurryScope to Scanimus, and the reason why we renamed it for obvious reasons. We don’t want to include the name of the primary defect that we’re trying to resolve in the name of a marketing strategy. So I thought Scanimus just sounded more elegant and more presentable at conferences and so on.

Michael Fanous: And so from BlurryScope to Scanimus, there must have been eight or nine different prototypes that we were playing around with, just little things, little edits in terms of the dimensions, things that You may appreciate subconsciously, but you would never say, “Oh, look, this is 35 centimeters tall, that the angle of the screen now is 20 degrees tilted a little bit.” But it was all done in a way to make it more appealing and more inviting for the average pathologist. And it’s very open, the system, and that took a while to figure out that would actually be little bit more appealing to a pathologist.

Grant Belgard: How do you decide which problems to solve with optics or mechanics and which to solve with computation?

Michael Fanous: Yeah, that is a really good question. So that is something that I wanna keep asking myself constantly, and that’s something I think we should Nowadays, especially with the power of AI, because so much can be reconfigured and reinvented on the mechanical end, I think we should constantly suspend our notions of what needs to be physically there and what cannot be challenged. We need to continuously challenge longstanding notions of like what is traditionally acceptable and what can be replaced or edited or reconfigured using AI. And I think the, the boundaries for this are limitless. It never ends, basically. You can keep trading. It’s a constant push and pull, playing around, editing, changing.

Michael Fanous: And if you take it all the way to the extreme of computation and you say “let’s just get rid of every single component and compensate for it with AI,” and then you’re left with one photodiode, and it’s just, you’re inferring everything from the most minute signal. That is, of course, not realistic. So you wanna be able to capture as much information, but what you’re trying to do is increase speed, reduce price, improve the portability. Those are the three main factors that you’re trying to optimize, and I think it never ends. I think with all of this new printing of any type of material really, sky’s the limit, and, your imagination, given what the AI could do and the algorithm and building around the AI, which is something that’s still novel.

Michael Fanous: Most people, when they’re trying to make so-called intelligent devices, they’re just tagging on at the tail end the algorithm and using the same data but then just processing it differently. But the real channel to true novelty and real advancement, in my opinion, will be doing the precisely the opposite, asking yourself, “the data can be processed any which way you like. Now let’s reconfigure, redesign the whole mechanics.” And so asking that question constantly I think is really healthy in this sphere and pretty exciting too.

Grant Belgard: How do you distinguish recovering information from generating something that just looks biologically plausible?

Michael Fanous: Yeah. So that’s really important in our case because some of these models are exceedingly good at presenting you something that is very pretty. so the, the big mistake here is assessing things qualitatively. So it’s quantitative versus qualitative. And for us, we’ve used every possible technical scientific metric to assess the acceptability of the image, like SSIM, PSNR. And so the other thing that we do is we use classifications to compare it against pathologist determinations. After that, we need to consult with a pathologist. Just asking yourself if it’s pretty, because some of these models especially, and you need to be careful which model you choose.

Michael Fanous: So with our model, we’re never reinterpreting features below a certain resolution which is important because certain models like Diffusion, for instance, they can be programmed to optimize for visible attractiveness and appearance and just how handsome is the image. And that’s not something you wanna be doing when you potentially may be compromising important medical features. So we’re very careful and cautious about that.

Grant Belgard: What evidence would persuade a skeptical pathologist to trust an output? Or what would you still tell them not to use it for?

Michael Fanous: Yeah. So if there’s any form of skepticism, I always like to compare with three different methods. So our system can handle all of this. They can do one, just the analog manual. You manually move the slide, and you’re looking at a live feed. That’s that’s one. The second thing you can do is you can have the traditional stop and stare method on our device. So you’re looking at just a crisp, sharp acquisition the whole way, and then the stitch is rendered normally. And then you can compare it with our deep learning deblurring system and see how you like it. Now if possible, and this pathologists typically have a set of slides that they keep at home or that they’re very intimately familiar with. They know all the details.

Michael Fanous: It’s like their precious little samples that they’ve just had for a number of years, and they know, and they’ve analyzed in different ways, and they wanna continue looking at it. So if they can bring the slide to me at my exhibit or on our device, that I think is the best way to convince them.

Grant Belgard: If scanning became much easier to access, what biological question would you want people to tackle?

Michael Fanous: Yeah, so I love that question because it’s a thought experiment. So if you take it to the absolute limit and you say to yourself, “Okay, now we’ve reached a point where digitization is trivial.” In fact, we have the luxury of digitizing more than we, we can handle. That’s not currently the case. In fact, currently, it’s the opposite. We have an unmet need. There’s a dire situation where we have too many scans that are not being processed and not enough pathologists. So it’s the opposite of what many people think. let’s say that we reach a point where anyone can digitize so easily that… and all of the classifications and the AI power and so on, so the pathologist now, all they need to do is make final approvals and make sure that with the edge cases and the extreme scenarios, that there are no problems. And then they have now, a very cushy career. It’s already good now.

Michael Fanous: It pays well and it’s a comfortable lifestyle, I think. But I think it’ll be really truly remarkable once all of the AI advancements come into effect. And then you have the option, I think, to reach for the stars, to try and do something absurdly scientifically novel and make some sort of breakthrough. And think about it, in terms of a biological insight that you would never find normally but now with AI empowerment, you can reach for a Nobel Prize or something. I know it sounds silly, but I think you can really strive for something very ambitious now at that point.

Grant Belgard: What do you try to learn from potential users and how has that feedback changed your priorities?

Michael Fanous: Absolutely. So when I go to conferences, when I’m meeting with pathologists and interacting with them, initially, I was just trying to ask myself, “Do they want this device?” That question was answered very robustly, very quickly. And so it wasn’t really a problem. Then I had to ask myself, ” what can I do to make sure that everybody likes it, and how can I please everyone?” But then I realized that’s the wrong question because if you have too many use cases, too many features, it gets too complicated, and that defeats the point. So the real question is, how do I make it irresistible to as many users as possible? That is maybe a cheeky question, but it’s the one I liked, and it’s the one that I’m still thinking about constantly. How do I make it irresistible? No excuses. We’ve pushed the mechanics and the optics and the AI as far as we can.

Michael Fanous: The ergonomics, the portability, It’s five pounds. When you lift it it’s unbelievable because this is a complicated optical device, and it just feels like maybe there’s a football in your hand. So that’s the thing that I’m trying to hone in on. How do I make it irresistible and make sure that there are no excuses for most people, that there are no excuses?

Grant Belgard: How are you Approaching funding, and what evidence do you think matters most to potential backers?

Michael Fanous: So that has been probably the biggest challenge on the entrepreneurial side, is trying to transmit the strong responses that I’m getting from pathologists to the average investor. The showing them and presenting the value proposition has been very tricky on my end. They’re mostly interested in figures and finances, and no amount of emails or, recommendations is really convincing for them. And I understand, that there’s quotas and they have to meet– they don’t necessarily understand the field of pathology. So that has been the trickiest aspect in terms of securing funding. It’s been very hard to try and transmit the response I’m getting at conferences and from pathologists, and to somehow distill it and bottle it up and then present it to them a way which they interpret as profitable.

Michael Fanous: that has been certainly very challenging and continues to be, I suspect, will be until this becomes widely adopted.

Grant Belgard: Which capabilities Are usable today and which are still research questions or future ambitions?

Michael Fanous: So we we have a lot of things that are in the works that we want to implement. Because it’s a research use only device, we’re very cautious with classifications, for instance. In terms of speed, we can always push the boundaries of that given how robust the AI is. So we’re still quite conservative, even though it doesn’t look like it. When I’m showing you the demo, people are saying, “It’s going so fast,” but they’re not saying it like in a negative way. They’re kinda saying, “And I think I may enjoy actually this whole new phenomenon here that I’m watching unfold before me.” So we can always go faster actually with the same frame rate, but using more robust networks. then there’s all sorts of different modalities that we would like to try and implement.

Michael Fanous: So there’s no end to the amount of features that can be included with this sort of scanning style, but we wanna be conservative right now. So I would say continuous scanning, deep learning, deblurring, that’s our niche. That’s what we’re trying to hone in on, and that’s what is our primary focus.

Grant Belgard: Where, if anywhere, would you connect this kind of imaging with genomic, functionalomic, or other molecular measurements?

Michael Fanous: Yeah, so that would involve fluorescence, and this is something that I am so eager to try and put together and implement because it really lends itself beautifully to, a GANscan-like modality. Unfortunately, fluorescence will have to wait. I’m just gonna have to be patient on this, bide my time. Fluorescence will involve because of long exposure times, it complicates things. But for me as, as an engineer, I already have all these… I don’t wanna give too much away, but there’s crazy ideas involving new kind of filter cubes, and I really can’t wait to unveil this. Probably, though, it won’t be for another year, probably a year and a half, because there’s just so much to do right now with the basic bright field compound transmission system. So yeah it’s definitely on our list, fluorescence with different types of fluorophores and so on.

Michael Fanous: And I think that motion blur would actually improve the situation and help you get more information. So that will hopefully be revealed 2028 or so. That’s– I’m very excited about that actually.

Grant Belgard: Let’s talk about how you got here. Before we get to degrees and job titles what were you like as a kid? What held your attention?

Michael Fanous: It may surprise people to know that I was very artistic I, I didn’t care for video games very much or computers. In, in the early 2000, I actually detested all things that were related to computers and software. I found it ugly and just unattractive and awkward. I liked instrumentation. I liked music. I was obsessed actually in my teenage years with guitar and violin, the possibility of improvising I thought was just so exhilarating. And I still have very artistic inclinations, but if I’m honest with myself, my aptitudes are more on the technical side. I was always better at math and physics especially was my best subject than I was at, for instance, literature or painting or anything like that, even with music. But I like to try and merge those two worlds together now. I think of a microscope as an instrument.

Michael Fanous: To me, what we’re building is one of the most exciting objects because of all of the degrees of freedom and all of the intelligence that’s involved. And so also, like histology data, to me, that’s like a painting. That’s like a work of art. So I see things in a very like romantic, artistic lens, and it used to be the arts and sciences. If you go to the Getty Center and you look at the microscope there, 1750s, half of it is a sculpture. It’s all filigree, gold, beautiful, intricate, ornate all of this stuff, and that’s half of the system. Because at that time, it was about being erudite and also fashionable, and those two things were not distinct. Now it’s like opposite. It’s diverged. Art and science couldn’t be more further apart, and it’s like the uglier the device, now you know that it truly has many functionalities, and it’s a very impressive advanced system.

Michael Fanous: where I think we should try to go back and try to make things a little bit more attractive, and I don’t know if that resonates with anyone, but that’s my kind of mindset.

Grant Belgard: How did you decide what to study, and what did that decision look like from the inside?

Michael Fanous: My father was very keen for me to be a doctor and work for him basically, and I really disliked that idea. I remember I thought I could never, work in hospitals because I’m too sensitive, and I immediately start commiserating with everybody. And I’m not trying to say, certain doctors are less sensitive or anything, but especially pathologists, they, I think can sympathize with this. They’re not really interacting with people directly. They’re interacting with slides. They’re alone. They can listen to music. They can be in their world, and so that’s something I have in common, I think, with pathol– And that was… If I was going to be a physician at all, it would be something like that. So I decided to pursue math, engineering because that’s what I was good at, very basically.

Michael Fanous: That may sound not terribly romantic but that’s what I was good at, so that’s just what I did because I knew I would be reasonably competent and I could pass, the exams okay. Because with physiology and biology, I remember I struggled terribly. I don’t have a good recollection. I have difficulty memorizing things, and that’s why being a physician or a doctor would never work out. There’s just too much to recall

Grant Belgard: How did You find your way to work uh, connecting biology, optics, and computing?

Michael Fanous: So I’ve always been fascinated by light. To me, light is magic especially on the visible spectrum, what is light? The more you study it, the more you realize that scientists are just approximating the different categories. It’s a ray, it’s a wave, it’s a beam, it’s a photon. Nowadays you’ll go to conferences, and it’ll be optics and photonics, as if these are two distinct things, it’s like optics is… No, optics is this. It’s the microscope with the incoherent illumination, and photonics, that’s, lasers and stimulated emission and all this, and I’m thinking, “Okay they’re really just two words that mean the same thing.” And it’s all just magical to me and extraordinary. With biology for me, the thing that’s most fascinating is human biology, the human body. That’s life. That’s organic, and there’s real intelligence. The brain is the original intelligence.

Michael Fanous: That is the authentic intelligence. So that to me will always be more remarkable to me to see intelligence in a human being than in a machine, and I think a lot of people will feel that way as well. And so combining what’s called the field of biophotonics or biology and photonics, that’s just a natural marriage, which is you don’t need to sell it to me. You don’t need to make the pitch. That to me is immediately very attractive. Computation, I would say is like the junction. It’s the connection between them. It’s the way it communicates. When you take an image of a biological specimen, you have now some data which can be processed using computation. So that’s… it’s the language that bridges them, I would say.

Grant Belgard: How Did you choose your doctoral research direction, and what were you hoping to understand?

Michael Fanous: So initially I joined niche kind of lab because it had computational imaging, basically. And it was called Quantitative Phase Imaging Lab. And extremely niche and the vision for my doctoral program I really didn’t have the capacity to pick the projects and make big, bold ideas. At that point, when I started my doctoral program, I just wasn’t there. There was a lot of work to do, I had too much to learn. So really, it was just trying different directives that my advisor gave me, and along the way, certain projects failed, and I faltered terribly. then other… And not necessarily what I expected that took the longest. Projects, they take two or three years, and it’s barely making it through the submission process and, it’s not a terribly impressive journal necessarily. And then some other ones, they go so quickly, and they immediately are accept– It’s just, it’s totally different.

Michael Fanous: So back, that’s when I reframed my whole thesis around pathology because those were the ones that did the best. And so I just ended up pursuing the combination of deep learning and pathology, and I was really happy with how that turned out. But it wasn’t by design and that wasn’t the inception.

Grant Belgard: What did your time at Illinois teach you about how to do science beyond the techniques?

Michael Fanous: But I should say that when I got to Illinois and when I got finally into the lab group, I felt so inferior to all my colleagues. They just surpassed me in every way, especially technically and, setting up experiments. And a lot of people don’t necessarily know this, but at the higher levels of academia, in physics and optics and these things, you need to be able to know a lot of information. It actually helps to have a photographic memory. These are the people that tend to excel the most, at first at least. And like, a colleague of mine, Chen Fei, he could do 3D Fourier transforms in his head. And for the first few years, I was trying to catch up with him and play that game. But it took me like three or four years to realize that I could never compete with this sort of style.

Michael Fanous: And also, the way that they picked their new projects, they would find the incremental progress in that niche field. I could never… I would be at meeting after meeting, and I always felt like something is off here. I’m not making any progress. And it took me finally, like the fourth or fifth year, I would– I stepped back and tried to think of things outside of the boundaries and the, the box that is our lab and our specific subjects. Because once you do that, then you have more possibilities, and you can always bring it back to the niche field that you are in. I’m a bit of a daydreamer. I get lost in distractions. And I always thought that was useless. But actually, some of my best ideas appeared ludicrous and silly and unworkable at the time. But when you tame it a little bit, and this is what I learnt ultimately. And I think I had to do this the hard way.

Michael Fanous: It had to be being humiliated and feeling inferior for a number of years until finally I just accepted I cannot win at this style and that game. I’ll never be able to compete at that level at specifically those faculties. I have to change the mindset, step back, look at the whole situation from a wider angle and be a little bit more imaginative, and then try to narrow it down and fit it within whatever topic and whatever resources are available. It’s just trying to reorient your whole perspective instead of just like learning things, which is also important. I learnt a great deal from my advisor at the time, who was very mathematically oriented. He was very physics-based, like theoretical. He was a very theoretical kind of guy, would derive all sorts of crazy things right in front of us which was inspiring, but it’s just not something that I could really aspire to.

Grant Belgard: How Did you find your way to UCLA and what were you hoping to get out of that?

Michael Fanous: That’s a very interesting story. Professor Ozcan, Aydogan Ozcan, a big professor at UCLA, his work was so popular in our specific field. I remember we would have meetings where we would discuss a whole new project and a whole new topic, and we thought we had found something novel and really exciting. And then we realized that Professor Ozcan had already published it two years ago. So this happened again and again. His productivity was like legend, and the, the caliber of his work was very high. And so I had been an admirer of his for, for some years. Out of nowhere, I realized one day ’cause near the end of my doctoral program, regrettably, my professor passed away and that was… I was sad about that, and I just abandoned a lot of my academic, ambitions.

Michael Fanous: And one day someone sends me an email saying, “Oh, you should look at this.” And Professor Ozkan had written a review on our paper, and I couldn’t believe it. so that’s when I reached out to him and I thought “maybe I could take this a step further. If he’s interested, it would be great.” I didn’t think, would, I would really have a chance. But once he wrote the review, I thought maybe there is, an opportunity here. And thankfully, he gave me a shot. I went to UCLA we took GANscan to the next level. It’s really been a remarkable journey.

Grant Belgard: Who has most influenced how you think about problems, and what do they do that has stayed with you?

Michael Fanous: My father actually has a lot of advice, and a lot of it he’s not a very technical guy so he doesn’t necessarily know all the engineering details and so on, but he’ll give me general bits of counsel. one that s- that sticks with me and I’ve always used is he tells me, “When you don’t know what to do nothing, and time will solve the problem.” And that has been to me proven to be very true and extremely helpful. Oftentimes I rush in to solve the problem. You feel this pressure and the need to solve something quickly and release the tension. But if you just stay with the tension and let time pass, often the answer will be revealed to you over time. And I think this can apply to anything almost. So I just wait. I actually let time solve pro-… time is an extraordinary collaborator for me. And I think it, it’s probably in general a, a useful tip.

Grant Belgard: Which part of the work has taken you furthest outside your existing training, and how did you go about learning that?

Michael Fanous: Definitely anything entrepreneurial is very much outside of my comfort zone, and for better or worse, I can’t really change who I am to suddenly become like a Type A person, very aggressive closer, and pitching, with, like, all of the, the right terminology and convincing, compelling phrasing and so on. That’s not my style. I am an engineer, technical person at heart, and that helps me in certain situations, for instance, like sales. I am quite, timid around the booth. I don’t aggressively try to pitch people and ask, pointed questions about how this could their particular workflow and then follow up aggressively. I’m incapable of that. I’m just having… I’m just a normal guy having conversations, trying to work through the problems, and I experimented for a while, know, with trying to play a different character, and that just didn’t work.

Michael Fanous: But, I think there’s still some room left to maybe refine areas and be a little bit more at the finances and at pitching, things like that.

Grant Belgard: Has An experiment ever changed Your mind about an idea you were excited about?

Michael Fanous: Oh, yeah. I, there’s so many instances where I’ll be sitting at a cafe, listening to Mozart or something, and sipping an iced tea and thinking, “Oh, I, I just stumbled across the most glamorous concept. It’s just, it’s gonna be beautiful. We’re gonna go to the lab. We’re gonna publish it in Nature. It’s gonna, e- everything is…” And then the next day, I’m in the lab, and it’s a complete catastrophe. Because empirically things don’t always work out like what you imagined. And so one particular example was, I can tell you, this is at U of I. I had this crazy notion of improving the temporal resolution of monitoring live cells ’cause there’s a problem with phototoxicity if you expose them to too much light. So I was trying to interpolate frames and then create this sort of glamorous video of a very slowed down high frame rate scene.

Michael Fanous: And I had all these crazy interpolation schemes and training scenarios. and the movies were, they were abysmal. And I remember my professor was trying to encourage me, but really the situation was pitiful. So after six or seven months we eventually abandoned it. But it was really one of those examples where I thought this was so promising, and this was really gonna be working out beautifully. But in practice, the reality was just the complete opposite.

Grant Belgard: How did you Make the decision to start Fanous Photonics and what felt most uncertain at that time?

Michael Fanous: GANscan had a commercial aspect about it, but it was all proof of principle. It was all theoretical. We didn’t– It would have been a licensing play and I didn’t really feel like pursuing that. By the time a device was actually built, again, that– this wasn’t the first thing on my mind at all, and it wasn’t by design. It wasn’t what I had been thinking when we put it together. But it was just nagging at me, there’s a device here. It’s a physical device, and I thought “maybe one of the undergrads will go and make a company and do something about it.” But, months would go by and no one was taking action, and I realized no one would take action. And I’m pretty much the one that is gonna have to do it if anyone does it at all. And I thought there was a real opportunity. I really didn’t know about, would pathologists really buy this? Are they really gonna be interested? Do they really care?

Michael Fanous: There was always that fear and insecurity lingering, lurking behind in the background. But There was also that strong sense that this is a device here. This is new. This is novel. Let’s– I could see this potentially working out. So there are two conflicting, notions and forces, and I finally gave in to just the one saying, “This is an opportunity. If you pass it now, it, it may not come again.” So that’s when I decided to follow through, take action to form the company. It was, again, it wasn’t like immediate. It was only really when I took it to a conference and I saw, okay, the demand here is real. This is tangible, because when you show it to family or friends, or even pathologists that you know, that’s not necessarily the right way to gauge interest because they may just be doing it, they may just be saying it to please you.

Michael Fanous: But when I realized that there was a real demand, that dissolved a lot of my insecurities.

Grant Belgard: Which habits from research have been helpful in building a company, and which have you had to unlearn?

Michael Fanous: That’s a tough one. Let’s see. I would say with research, being extremely flexible and always changing your mind a good thing up until the end, up until the final submission. even the title, the whole value of the work can be re-edited. is not necessarily the mindset you wanna have when you’re pitching to investors, when you’re going about trying to sell the device. You need to have a concise message which is same and unaltered and in cement and crystallized. That way it’s easy for people to digest. You can’t be changing your… nobody knows what you’re talking about. They think you’re selling one thing, and then you’re talking about something else, and there are too many features, and it’s all just blending in this bizarre amalgam that is not helpful. That’s something I need to work on still, I think, ’cause I like playing around and continually changing things.

Michael Fanous: But in terms of marketing, this is the last thing you want. You wanna be precise, consistent, re-repetition.

Grant Belgard: A life scientist who wants to work at this intersection of imaging and machine learning, what would you suggest they spend their next three months doing?

Michael Fanous: My advice, take it or leave it, but if you wanna make the biggest impact, I think look at the hardware, study the hardware, because that is really where there’s the most possibility and the greatest chance for truly something groundbreaking. If you’re just playing around with the software and the various AI tricks and you’re, using LLMs, that has a limitation. There’s only so much you can do. But if you examine the hardware and think how can this be reinvented?” And you only need one great idea. So you can spend three months failing, repeatedly, thinking of all the wrong… And the LLMs won’t feed this to you. This is something they still struggle at, getting really original, truly groundbreaking ideas. You can feed it all you like and play around with the deep search mode or whatever mode you like. I’m still playing with that, and it’s not giving me something truly unique.

Michael Fanous: So I would say this is something that’s still in the domain of human dominance. if you’re gonna pursue something for three months, you really only need one great concept a piece of insight. And if you focus on the hardware, I think that’s where there’s You know, you can, you could revolutionize something with just a small tweak in the hardware. That’s where I would say that’s a goldmine right now. So focus on that. That’s my take it or leave it. That would be my advice.

Grant Belgard: How should a scientist de-risk before deciding that a research project deserves a company?

Michael Fanous: Yeah. that’s, again, that’s a tricky one. For me, it was an, an increasingly loud voice nagging at me that this really should… something should be done about it. I think if there’s a lot of hesitancy and it’s not in the hardware space. Again, this is counterintuitive and contrary to what most people would recommend nowadays especially, ’cause I’m pitching to investors and when I say this is hardware-based or primarily, that’s a red flag. They dismiss it. They’re thinking this is the AI bubble. This is the era of software. And I’m thinking no, no. software players, that is very dangerous right now. You do not wanna be just playing around with code. The LLMs are gonna eat you up and in the next few months. So that’s what I would caution. Again, this is not necessarily a popular opinion. This is just coming from me and from what I’m seeing and my experiences.

Michael Fanous: don’t feel like software is safe. Hardware, and especially novel thing that are designed specifically around a new AI concept, that’s if that’s what you have, and you have it, and it’s truly novel, and then you make sure that there’s a demand, I would say would be something to worth worth pursuing.

Grant Belgard: What advice would you give your younger self at the beginning of this journey, and would your younger self have listened?

Michael Fanous: Wow. So I probably would not have listened to myself because I typically don’t take advice. There’s too many conflicting pieces of counsel that are coming my way, and I’ve made a decision early on to just try and a lot of the… But if I could get through to me, to my earlier self, I would say trust your intuition more because a lot of people disagreed with many micro decisions that I was just doing naturally because I intuited. I just– my intuition dictated. No, There has to be a screen here. It needs to be tilted. A lot of things like that, that no… And also the demographic. People were telling me, “You need to start in developing nations,” and things like… and my intuition is just screaming at me no. The West is the problem. We have serious, dire, unmet needs.

Michael Fanous: The technology in the state-of-the-art clinics is Like, when you compare that with what people are showing at these conferences, and they have their badges, and they’re sh– and they’re telling you all of these things that are so extraordinary. And then you go to the average clinic, and you realize, wait a second, there’s a time difference here of a hundred years. So I would just say I would amplify that intuition and be a little bit shameless about listening to the things that are being said in, in the background of my mind and not necessarily being too scared about carrying that out. Because I was always so uncertain, I think if I could just tell myself, just let the intuition run wild. Just let it, just let it have its say, that’s what I would say.

Grant Belgard: What would you most like listeners to take away from this conversation, and where can they follow your work?

Michael Fanous: Yeah, so I hope that they en-enjoyed it, that it brought them some degree of amusement. I think, if it was any degree entertaining, I think that’s worthwhile. If there’s any engineers or any technical like AI kind of people that are listening, I hope that they adopt some of these contrary concepts of, starting with the AI being disruptive on the hardware side. I hope that they use that ’cause that– I really think there’s gonna be a great wave of innovation and novelty, and it’s really an exciting time. as far as anyone who’s more on, on the uh, clinical side or pathology especially, I, obviously, I hope that they are piqued and interested in the device. And will maybe ask for a demo and seek to come to the conferences or look me up. And you can follow me on LinkedIn. My name is just Michael John Fanous.

Michael Fanous: We also have a company page on, on LinkedIn, Fanous Photonics, and the website is scanimus.ai. We post all of our news and new articles that we’re doing, so there’s more publications that are coming out. So that’s where you could follow me, and I’d appreciate it.

Grant Belgard: Michael thank you for joining us and sharing both the science and the story behind it.

Michael Fanous: Hey, Grant. Thanks so much for the opportunity. Been a real pleasure

The Bioinformatics CRO Podcast

Episode 93 with Trevor Nicks

Trevor Nicks, founder and CEO of Caravel Bio, discusses biological compute, protein spores, and how to leverage intrinsic biological properties for protein engineering.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Trevor Nicks

Trevor Nicks is the founder and CEO of Caravel Bio, which innovates using cell-free protein engineering.

Transcript of Episode 93: Trevor Nicks

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. Today, we’re speaking with Trevor Nicks, founder and CEO of Caravel Bio. Trevor is a biotechnology engineer and entrepreneur whose path has included an early algae biotechnology venture and doctoral research in chemical and biological engineering at Tufts, where he worked on bacterial spore display and protein stability for cell-free systems. Our main topic today is biological compute. One of machine learning’s early lineages, genetic algorithms, borrowed a simple loop from evolution: vary, test, select, and repeat. Trevor’s work asks what happens when that loop is not merely represented in software but made physical, using real evolutionary cycles to search for better proteins.

Grant Belgard: We’ll explore how that paradigm differs from conventional protein engineering, where machine learning fits, how it might translate into applications such as critical minerals processing and industrial separations, how Trevor arrived at this work, and the advice he would give scientists and founders pursuing unconventional paths. Trevor, welcome to the show.

Trevor Nicks: Thanks, Grant. Glad to be here

Grant Belgard: So for listeners meeting you for the first time, what are you building now and what problem are you trying to solve?

Trevor Nicks: So I’m building Caravel as a company, right? And within that Caravel is a platform, and the problem our platform solves is how do we generate data that will allow us to actually create products that work in the real world? And so we ask really important questions like how do we generate more data? How do we generate data that’s relevant to the final application and not just in proxy conditions? And how do we generate that data in a timely manner and with costs that are low enough that we can actually build products before investors or customers get impatient with us?

Grant Belgard: So what do you mean by a protein machine, and what does that phrase reveal that protein engineering does not?

Trevor Nicks: Good question. A protein machine in the way we think about it is that there are simple machines, which might be a singular protein and then there are more complex machines that might be multiple proteins or protein systems. And so at Caravel, as you said in the intro we leverage a unique piece of biology called bacterial spores. And spores have this protein shell, and then we’re engineering that shell as a carrier that can be functionalized with many different proteins or enzymes. And so when we say protein machine, we’re often talking about along this layer of singular proteins, which would be enzymes or therapeutics, all the way up to these larger protein masses that have many different pieces of protein machines upon them that are a more complex machine that can then serve a, a certain function.

Grant Belgard: When you use the phrase biological compute, what do you mean?

Trevor Nicks: Compute for us is often thinking about the idea that we need to have memory. If we’re trying to solve a problem, we need to ask how well did it go last time?” And we need to have adaptability, so thinking about what are we gonna do for the next round when we try to solve the solution in, round two or three or four or five or a hundred. And then we need to have the ability to move between those rounds very quickly. And so biological compute here for us is thinking about how are we doing evolution, what systems, like biological systems, are we using to do our memory and our adaptation but then also the layer of machine learning and computation that goes alongside our biology. So we’re not just doing biology, and we’re not just doing machine learning. It is a hybridized system where there’s a whole loop that plays together, interacts round to round.

Grant Belgard: Can you narrate one complete mutate, test, select, repeat cycle as though we were watching it happen at the bench?

Trevor Nicks: I’ll give you an example from our work on carbon capture with Shell. As I was just– as I was leaving my PhD, we started to have a good conversation with Shell through the Greentown Labs program on how could we make enzymes enable us to do carbon capture better as a society. And so there’s this idea that people have been working on for about twenty years that enzymes could allow carbon capture to be lower cost by changing the solvent that’s actually doing the capture. And an enzyme could allow us to use a lower cost solvent that requires less heat to pull the CO2 out in the end. But the same type of chemistry change that allows us to use less heat and save cost means it also absorbs CO2 slower. So could we use an enzyme to speed that up?

Trevor Nicks: People have been trying to do that for a while and there are some issues there around the industrial viability of the enzymes to actually do that job in a costly or cost-effective manner. And so the loop that we’ve created for that has been asking this question of how can we make an enzyme that lasts as long as possible in these real industrial conditions? And so because we’re using our spores, what we do is we take the spores, which have a copy of the genome inside of them, and they’re expressing an enzyme on their surface. And we make, a million different versions of those spores, and we can screen them in the actual industrial conditions because of the stability of the spores themselves. And when I say conditions, people might be thinking pH and temperature, and that’s true. But on top of that too, as importantly, is time.

Trevor Nicks: And the true phenotype that we want from these enzymes is how long do they last in a given time cycle. For this exact experiment, we’re actually relying on the ability of the spores to maintain a specific genotype-phenotype link for months on end, and so that we can ask this copy of this enzyme, how long did it last in the exact industrial conditions? And so that loop is make spores, put spores in reactor run reactor for two months. Afterwards, pull out all of those spores and ask which ones are actually still viable. And then from that, learn and repeat. So it’s a very long loop because the goal is longevity of the enzyme itself. That’s what we’re evolving for.

Grant Belgard: What sense is that cycle computing rather than iterative screening?

Trevor Nicks: The idea of compute is the memory itself, in that in traditional systems for doing approach in engineering, it would be essentially impossible to ask for a million different versions of these enzymes. How long do they last in these scaled conditions in the reactor? Because there’d be no way to maintain this genotype-phenotype link in a way that is manageable or viable. And because we’re able to maintain that memory over that long period of time, that actually allows us to do the measurement that enables the compute to then get to the next round of predictions whether that’s through random mutagenesis with a genetic algorithm or if it’s through other types of algorithms.

Grant Belgard: What parts of the search are best handled by evolution, which by machine learning and which by human judgment

Trevor Nicks: That’s a fun question. I think we are actively evaluating that for ourselves right now. I don’t know that I have a perfect answer for you. Right now, in the carbon capture work, we’ve had some really good luck with really early designs that came out of building models based off of publicly available data. And then some of our early mutation data baked into that gave us some really great enzymes just in the first six to twelve months of this work. So that was definitely a hybridized system. In other areas where we’re working, for example, in the critical minerals, we’ve tried some work to do de novo protein design for the minerals, and there’s been some luck there.

Trevor Nicks: But some of the best designs have actual- actually just come out of the fact that we have someone who has a PhD in metalloproteins, and his instincts on what the chemistry should be– And by instincts I mean he studied at UC Davis and UC Berkeley, for ten years and is very good at understanding these chemistries and how the different shells and amino acids are going to interact. And that’s given us actually better proteins thus far than any singular machine learning prediction sequence. But we are doing the combination of both and seeing if we can, in the end, create a system that enhances the abilities of people to design better proteins through physics-based modeling.

Grant Belgard: How do you design a selection pressure that rewards the property you ultimately need outside the assay?

Trevor Nicks: For our system, we love generating lots of data. And so if we can, we love to come back to what is called fluorescence-activated cell sorting or fluorescence-activated droplet sorting. And so within that is we are trying to do our best to, again, maintain the memory of what occurred in that industrial condition because the spore is going to maintain whether or not this enzyme or protein was inactivated or damaged during some given process, and then evaluate if that protein is still viable while it’s being sorted, whether that’s in a microfluidic device or in a laser-based system like FACS.

Grant Belgard: How do you distinguish transferable improvements from assay artifacts?

Trevor Nicks: Yeah, that’s a good one. We’re not always able to, right? And so it– when I say that, I mean in that we’re not always able to immediately, and then it’s oh, we learn when we want to scale up that oh, that was just the thing where they performed better in the assay. So for example, in some of our carbon capture work we did have enzymes that were effectively false positives, where they would sit in the reactor for a long period of time, and then afterwards we’d evaluate them. And there were some that were active when we were actually evaluating them after that two months. But then when we actually tested them in bulk where it’s like that pure enzyme, it didn’t actually work at the elevated temperatures and with like the other contaminants that go into the actual industrial reactions. And so we effectively learned oh, that enzyme had just refolded when we went back to test it afterwards.

Trevor Nicks: I think the answer to your question is, we always do our best to approximate the real world conditions and measure in those but there has to be a cycle of veracity where we’re funneling down. And every time we funnel, we are asking, “How does this actually work in the scaled conditions?” And we try to get to that point as fast as possible.

Grant Belgard: How do you keep a search from losing diversity or settling too early on a local optimum?

Trevor Nicks: We do a couple different things there. The first is At a high level between rounds of evolution, we’re always doing mutation. So even if we had, from our last round oh, just this one sequence dominates the population. It’s ninety-nine percent of what was enriched. We’ll take that one, and we will add back in some of the ones from the previous round while also recombobulating those together just to make sure that, we’re not just gonna get that same one again after the next round of selection. And then at the same time, we’re diversifying from that one. Now, diversifying from just the one that does mean you’re probably gonna be around that local maximum.

Trevor Nicks: But then because we’ll also recombobulate it with other ones, whether those came from nature or if those come from AI designs like de novo predictions that recombobulation on its face is giving us a broader search space than if we were just to do the evolution itself of always saying, “Okay, this is the parent. We diversify from the parent.” Because we always seed back in a few other designs, we’re getting a more expansive search space, and that’s very purposeful

Grant Belgard: Where, if anywhere, do cell-free systems change what can be explored?

Trevor Nicks: I originally really wanted to do my PhD on cell-free protein synthesis ’cause it was so cool when I was, looking at grad school that like, oh, you could express these toxic proteins or maybe you could put the, put in a non-canonical amino acid super easily. And that was really exciting. Didn’t get into any labs that were directly doing cell-free protein synthesis at the time but ended up working on this spore project which in the end has actually enabled us to do cell-free protein synthesis in really new, powerful ways by increasing throughput. And the point being that we’re using cell-free protein synthesis today to make proteins that are toxic to living cells. That’s kinda the first thing. So we’re engineering this protein where if you overexpress it in pretty much any bacteria or even mammalian cells, it kills them because it’s damaging the genome.

Trevor Nicks: But we really need that protein to work ’cause if it does work, it’s gonna allow everybody in all of biotechnology to have access to lower-cost DNA. And the fact that it enables lower-cost DNA is also why it’s toxic to express in cells, right? And that’s an example. The other thing here is I’ll circle back to, again, genetic code expansion. So that’s where we’re making systems that use amino acids that aren’t normally found in nature. To do genetic code expansion inside of a living cell is quite arduous in that it can take a lot of time to do the background work of doing the genome engineering to actually enable the use of that amino acid at a high enough level for you to actually be able to study it or use it. We are able to do cell-free protein synthesis to quickly prototype and ask the question, is it actually worth using that non-canonical amino acid in this protein?

Trevor Nicks: Does it actually make a good product? And then if it does, if the answer is yes to that question, then we can put the time and money and invest in building the strains that can make that at scale. So that’s kinda how we’re thinking about it today. It’s either making things that we couldn’t make in a cell or making things much, much faster to then actually see what we should invest in to make the full product.

Grant Belgard: What role in practice does genetic code expansion play in expanding the searchable chemistry space?

Trevor Nicks: I think it has a huge role to play, but importantly, it’s not the only thing that has a role to play. There are other systems too. But so for genetic code expansion, there are over five hundred known amino acids in the literature that people have made systems for. There are amino acids that are halogenated or have new kinds of click chemistries, which is like how people are making antibody drug conjugates. And so we’re, looking at, all those different kinds of amino acids that other people have developed and worked on, and also thinking about some new ones that can potentially enable some types of um, new chemistries that don’t exist yet in biology. But then beyond that um, other types of new chemistries, or I should say new chemistry-enabling systems or like some types of proteins that enable post-translational modifications.

Trevor Nicks: Our chief innovation officer Agneya, he did his PhD and postdoc at ETH Zurich studying these microbes that live in sponges in the ocean. And they have this entirely different system that isn’t found elsewhere for making post-translational modifications in a very targeted, specific way. And so we’re also looking at those systems to ask, can we use those to make proteins that have new chemistries? And would that actually enable those new proteins to scale better than if we were using traditional genetic code expansion? So we’re really excited about genetic code expansion, but even more importantly, we’re just asking how do we make the best chemistry? How do we make the best protein machine, if you will? And then checking out what tools exist to make that happen.

Grant Belgard: Hardest about maintaining a reliable link between sequence and measured function at very high throughput

Trevor Nicks: I think the hardest thing is one of course is like the assay to even have the throughput. But then it comes back to the question you asked earlier in terms of how do you distinguish between an artifact of your assay and what’s actually useful. And to give you a bit of an example, sometimes we’ve made enzymes that in theory, based off our ultra-high throughput assay, were extremely active and doing much, much better than any parent enzyme, like fifty X. And then we actually went and scaled that up and we learned oh, that wasn’t even relevant because the actual substrate concentration is so much higher in the scale system that change in speed didn’t even matter, right? So I would say the hardest thing about ultra-high throughput screening is making it matter in terms of having the data you produce, the new sequences you produce be relevant to the final product.

Trevor Nicks: And that’s a learning curve for almost every new protein. And so making systems that allow us to run that learning curve faster for new proteins is a big part of what Caravel is becoming in terms of making it applicable across many different protein classes.

Grant Belgard: What makes the resulting data especially useful or especially difficult for machine learning?

Trevor Nicks: I think I’m gonna go with useful on this and it’s volume of data. If you can make this stuff work where you do get an ultra-high throughput screen and you are generating, potentially hundreds of millions of sequence function relationships, you can make some really cool models that don’t require protein structures to be useful. And protein structures are very slow and arduous and expensive to make, generally speaking. And so if you can generate enough sequence function data that is relevant to your final product that can allow machine learning to do some really cool things.

Trevor Nicks: And so while we’re in the early days of really proving that out across multiple protein classes, we’re excited for that possibility in terms of creating really enormous data sets across many different protein classes and many different, again, like scales of where protein machines can be used in terms of testing them in the scaled environment that we actually want to use at the end of the day to then train models to not just, think about new proteins, but actually new products. In terms of protein engineering on its nose is a multi-parameter optimization problem, and we don’t need a model that just improves thermostability. We also need a model that lets us know how manufacturable is it gonna be. Does it also tolerate the pH? Can it be spray dried? All these other questions that are really important, and we don’t really have a PDB version of at this moment in time

Grant Belgard: How do you keep a model from learning the quirks of an assay rather than the underlying biology?

Trevor Nicks: I think the way we approach that is having more than one assay if possible. We, for example, with some of our carbonic anhydrase work for the carbon capture, we are working on multiple versions of the assay at ultra-high throughput, some of the forward direction, some of them reverse direction. And then having two versions of the forward, so that if you do have a effect that’s mediated by the fluorescent reporter itself, then you can mitigate that effect by going back and forth between the two different versions of the assay. That idea itself was actually something that was first suggested to us by one of our advisors, Kevin Gray, who had done some enzyme engineering for amylases back in the early 2000s. And they had this really fantastic amylase that they were really excited about. And when they went to scale it up, it was like, “It doesn’t work at all.

Trevor Nicks: What the heck?” And it was, like, because in their assay, they had this other molecule in there that was essential for the assay, and it turned out the enzyme had effectively evolved to use that as a cofactor, which is crazy to think about. But that’s what biology did. It found the best system with all the tools at its disposal, and we went to scale it up, and that cofactor wasn’t there anymore. You know, it’s crazy what’s actually happening in these systems. And yeah, having more than one assay is kinda how we think about doing that.

Grant Belgard: How do you think about digital compute and biological compute? Are these substitutes, complements, layers of one system?

Trevor Nicks: I think absolutely complementary. We are thinking about this from the perspective of machine learning is useful when you have data, and it’s useful when you have data that is relevant, again, to the actual product you want to make, right? And so at a high level, our goal is to do artificial evolution effectively in the lab in these new systems with our spores and with our temporary synthetic cell technology that enables the cell-free protein synthesis. But the Layer on top of that, of feeding that data into these machine learning models and seeing how useful they can become is, I think in my opinion, a good use of time and resources for us in the company. ‘Cause while it may not make the first product we build better, it could make our tenth product better, or it could make it faster to build our tenth product, right?

Trevor Nicks: So that’s what we’re working towards in terms of making sure we have all the infrastructure, that the data we generate today is useful to us two years from now, five years from now, 10 years from now when we’re building products in the future

Grant Belgard: Think us through one problem in critical minerals processing or industrial separations where this approach could be useful

Trevor Nicks: Yeah, sure. So it’s in the name, right? Critical minerals processing and separations, right? In separations, if you’re gonna use biology to separate something, you’re probably using a protein to do that binding of a given object, in this instance, the ion of a critical mineral. And proteins themselves are tiny, and separating proteins out from a liquid solution is a really hard problem. There are entire industries built on that, right? Like protein A for antibody production. And so if you just had a protein that bound the metal, that’s not a product because to actually do that at industrial scale, you need to then be able to pull out that metal protein complex and separate it out from everything else, release that metal, and then purify the metal, but then importantly, reuse the protein.

Trevor Nicks: Otherwise, the protein is way, way too expensive to actually use for that initial function of bind metal, release metal. And so from our perspective, we’re actually generating systems where we’re evolving these proteins in a format in which they are easier to separate and easier to reuse from the rest of the solution because we actually use the spores in our final process. So a quick microbiology lesson, right? Again, these spores are effectively bacterial cocoons that are extremely stable. They can survive temperatures over a hundred degrees Celsius, extremes in pH, solvents, et cetera. So they’re actually like a carrier for these proteins in these industrial extremes.

Trevor Nicks: And because we’re evolving our proteins on those carriers, we can be more confident that when we engineer them to bind a specific metal, that we’ll actually have that function at scale because we were already evolved on the carrier. Whereas traditional approaches evolve proteins in one condition, for example, in a yeast cell or in an E. coli cell, and then they spend, months to years figuring out, okay, now how do we immobilize this protein on a polymer so that it’s in this reusable format? So we like to say that we do contextual evolution. We’re evolving the protein in the context in which it will be used at scale. And we’re seeing that is really important in terms of having this transferability across scales and actually creating products.

Grant Belgard: How do you decide whether an opportunity should become a platform capacity, a product, or a partnership?

Trevor Nicks: At a high level, comes down to money. Who’s going to pay for it? How many people will pay us for it? Will it be a functionality that is useful for many products or just this one product? Is it even worth building for this one product? Generally speaking if we think of a new capability while planning an experiment that increases, for example, the number of genes that we’re able to get into our cell, we know that’s gonna be useful for everything we do in the future because it means we’ll be able to test more genes for everyone, for every client that we ever have or for any product we ever want to build for ourselves. That’s something that is definitely like a platform feature. But if it’s, for example an assay development thing where it’s like, we just have to have this version of this assay for this one product.

Trevor Nicks: And it’s gonna take us, five million dollars to build up that assay, that hyperbole. We might not do that. Maybe we just drop that product, right? And don’t pursue it because it’s just gonna be too complicated and it’s too much of a upfront investment before you even know if anything’s going to work to warrant it. So that’s kinda how we think about it in terms of this is gonna improve the platform for every experiment we ever do in the future, or this is a product-specific capability that doesn’t transfer, and if it’s too expensive to develop it, then we maybe just won’t even develop it

Grant Belgard: How do you get a durable advantage when the tools, data, assays, and products are all changing so quickly?

Trevor Nicks: One is humans. I think we, at Caravel, we’re placing a big emphasis on our relationships that we’re building with our customers. The second is we are fundamentally using a new technology that hasn’t existed before, that has value propositions that haven’t existed before, or at least not at this throughput. And so the amount of data that we’re generating itself is a moat. And as we continue to build up that moat and it gets wider and wider as we continue to generate millions of data points per day it, again, this is a hypothesis. Five years from now, we should be able to design a product faster and with less total cost than a competitor that may be starting out afresh with the new technology. So that’s how we’re thinking about that.

Trevor Nicks: One, establish and maintain relationships, and two, generate as much data as we can so that we can always be making our products and our product production cycle faster and less costly.

Grant Belgard: What safeguards will matter most if the searchable protein space expands substantially?

Trevor Nicks: At a high level, I would say that someone being able to produce a protein in a lab somewhere at a nanogram, you know, in a one-off assay because they write an academic paper that was publicly available, that’s one thing. That’s not, something I think of high concern. It’s more about someone being able to maybe potentially manufacture something that is dangerous, right? It’d actually be a concern for a large part of the population more than to the danger of the group in that lab. And so I would be thinking that, regulations around customer verification for DNA that’s going to manufacturing facilities or products that are being ordered that enable the manufacturing of bioproducts, maybe with expanded chemistries or with different types of new biology, of course.

Trevor Nicks: That’s something that should be considered potentially to be regulated in the future, and I think there are actively conversations going around that are pretty responsible about that in the administration right now as it relates to AI-designed proteins. So that’s kinda how I would view that, is that, we don’t want to hamper the ability of any American company or any company to make a biologic that heals people, right? Or helps people or makes food cheaper, right? But at the same time, yeah, of course, there’s safety concerns. But, smart, logical Controls around the supply chain itself I think makes sense in terms of controlling what can and cannot actually impact the population in a negative way.

Grant Belgard: What could count as a decisive proof point for your platform over the next one to two years?

Trevor Nicks: Scale. I think, for the past twenty years, a lot of biotechnology has operated out of the same organisms and, from the same production systems in that, okay you’ve made this protein in a lab somewhere, now you’re gonna produce that at scale, and you’re gonna do that in CHO cells or yeast cells or E. coli cells or, the past decade, different types of yeast. For us to say, “Hey, this is an entirely new production system too,” it’s a really big bar for us to get across in terms of… And yes, by the way, this does actually work at scale. The core thesis of the company is that we’ll make more scalable products. We actually have to make a scaled product first. So we’re excited and very thankful that we have customers that have bought into this vision and we’re really grateful to be working with them.

Trevor Nicks: Yeah, hopefully in the next eighteen to twenty-four months, we’ll have some things going at really large scales, and we’ll prove out that thesis.

Grant Belgard: Before we leave the word compute I want to spend a few minutes on the physical resources behind different computational paradigms. What, if anything, does Anthropic illustrate about the relationship between computational capability and resource demand?

Trevor Nicks: With AI it comes around to data centers and supply chains. I think we’re seeing this ourselves in the past two years, Anthropic or anyone, right? In terms of the, when the OpenClaw model became so popular and people started buying up the Mac Minis, and during COVID, when supply chains became constrained for rare earth elements and we couldn’t get chips and, semiconductors that require those rare earth elements for many different reasons. Because demand is increasing so much for compute itself, and compute requires and effectively eats electricity and also metals to be able to do compute, I think we’re seeing a huge increase in also now price and concern at a government level even around how do we make enough energy and how do we secure enough metals to satisfy this new demand? And so that for us has led to really two of our first products.

Trevor Nicks: On the energy side of things it’s how do we generate enough energy and bring up enough energy online without also endangering ourselves through excess natural gas release. Some of these new data centers are building really big, some of the biggest ever natural gas plants, and there’s a lot of CO2. Beyond that is the metal supply chains themselves. If we onshore the same technologies that were used for critical mineral processing thirty, forty years ago, those use a lot of chemicals that aren’t great. And not just to make an environmental argument, they’re also really bad processes . They are so slow, and they– if the factory turns off for a little bit, it can take six months to turn it back on. That’s not a good supply chain. That is such a bottleneck.

Trevor Nicks: And we’re really thinking about this from some of the first products we’re building at Caravel, not intentionally to begin with, but it just so happened that it’s lined up in this way that’s oh, the first two products we have, carbon capture and critical minerals recycling and purification, are two of the biggest things that AI eats, if you will. And so that’s how we’re thinking about it in terms of the physical commodities that are required to power these systems that everybody is using and is powering a lot of economic growth.

Grant Belgard: Let’s talk about how you to where you are now. What did you imagine your career would look like before taking this direction?

Trevor Nicks: So I grew up on a farm in Missouri. My older brother’s a pastor. My younger brother sells tractors. I went to college on a track scholarship, and I wanted to be a chemistry teacher ’cause I had an amazing chemistry teacher, among many other amazing teachers and thought her job was so cool. But then in college within my first semester, kind of, Oh wow, the world’s so much bigger than I thought. There’s so many other things.” And I was student teaching actually, and in my student teaching, I was shadowing an instructor, and she kinda just turned to me one day, and she’s like, Trevor, you’re really good at this. You could do other things too if you wanted that would make you a lot more money than being a high school teacher in Missouri.” And I think it was really good advice, and I’m very thankful for her for saying that.

Trevor Nicks: Not that I’m making a lot of money right now you know, running this startup, maybe in the future. But the point being that it has been less of a money-oriented journey for me and more of about exploration and curiosity and just excitement of building something and building something that I hope will matter in the future. And so yeah, I just turned that curiosity for chemistry and also the desire to share and talk to people. I would candidly say that being a CEO of a company and raising money is not that different from teaching in terms of I’m effectively teaching people my vision of the future, and this is how we’re going to build it, and I think you should support it in these ways. And that requires a lot of education when you’re talking about a brand-new technology that’s never existed before and is building and uniting multiple unique disciplines that are coming together.

Grant Belgard: Which early experience most changed your understanding of how biotechnology companies succeed or fail?

Trevor Nicks: I’ll go back to the summer of twenty sixteen. I had just signed on to be an intern for this company that had just gotten their first check from IndieBio EU at the time, which is located in Cork, Ireland. And the founding CSO, like the week or two before we were gonna leave the US to be in Ireland for the summer, decided he didn’t want to do it anymore. And so I got promoted from intern to chief science officer. You mentioned that in the beginning in terms of this early experience with the algae company. I was twenty years old. I hadn’t even taken a microbiology class. I had no idea what I was doing. But within that, there’s this cohort of thirteen companies that have all descended on Cork, Ireland for the summer to learn from some of the folks at SOSV and IndieBio. And I think the, like the most formative experience for me in that…

Trevor Nicks: The whole experience was formative of course, but like the thing that stood out and the thing that shaped my trajectory afterwards was like a singular workshop on techno-economic modeling. In that I realized, it’s like, “Oh crap, this algae stuff is never going to work at scale. The economics just don’t make sense.” And so that was for a little bit me and the team were like no, but if we do it in this way and if we get this, one version to be fifteen X better, then yes, it could be good enough maybe.” And then the company ended up pivoting to work on other things to be produced in the algae, which was really exciting for a time.

Trevor Nicks: I ended up leaving and going back to undergrad and finishing my undergraduate degree in biochemistry with the goal of being like I want to work on a system where fundamentally I think that the economics will be more scalable because I wanted what I was doing at the bench to matter in the end. And so that was why I started working on this system and was excited to to join the lab that I ended up joining for grad school.

Grant Belgard: How did your doctoral work become a foundation for what you later built?

Trevor Nicks: As you said in the introduction my doctoral work was on this grant that was funded by the Department of Energy to study and make new tools for spore display. And so spore display fundamentally again has this idea of engineering bacterial spores, which are these cocoons, to be functionalizable with almost any other protein. And so we built tools for functionalizing the spores so that when the bacteria makes spores, we can put many different proteins on them. So that was my PhD research, was like, how do you make these self-assembling, living yet dormant beads? That was really fantastic in terms of learning a lot about this organism that, a lot of people work in bacillus, don’t get me wrong, but not a lot of people work on spores themselves. So it was really fun to kinda like work in this niche area.

Trevor Nicks: And then beyond that that biology has actually become the foundation for our entire company in terms of how we do research. So Caravel itself is not a spore display company but we leverage the biology of spores in new ways. So the utility of the spores is much higher than I first appreciated when I was applying to grad schools. Effectively, what’s happening with the spores and the way the spores work is, again it’s a bead in that you can make a protein, and then it’s on the spore surface, and there’s no functional protein there except for the ones you wanted. And so effectively, it’s a really simple way to purify and study a protein, and that’s really powerful. If you talk to any biologist, purification of proteins is like one of the hardest problems that we face working at the bench. Keeps you up at night, keep you awake for months in terms of does it actually work?

Trevor Nicks: And so that ability to have this like self-assembling, genetically encoded bead is actually really powerful. And so now we’ve built that into new systems that leverage spores and microfluidics to enable us to study and engineer in ultra-high throughput almost any protein. I’m not gonna say we can do everything, but we’re doing a lot in ways that I never thought would be possible. And it’s an exciting system and it’s leading to, again, more data, more types of data that costs less, and data that’s more scalable because we’re uniting this like bead technology with cell-free protein synthesis that allows us to be very versatile and expansive in terms of what proteins we’re actually able to build and study as it relates to genetic code expansion, post-translational modifications using a human cell lysate or a tobacco plant cell lysate.

Trevor Nicks: So it’s a really fun system to be working with, and at the end of the day, it’s just easier for us to actually isolate the thing to collect the data in the first place. That’s why it’s useful.

Grant Belgard: What parts of translating research into a viable company was most underappreciated from inside academia?

Trevor Nicks: I think there might be more similarities than people give it credit for in terms of when you’re in academia you kinda have two customers. Your customer is your department head or your dean in terms of making them happy by doing your course load and everything. And then your other customer is probably the government in terms of getting your contracts right and getting grants to support your research. On the industry side of things, you also effectively just have two customers or two clients. And our clients are investors who giving us our pre-seed funding and we’ll be raising more money from in the future probably. And then also again, the government and customers who are giving us money to do work for them and to develop products. And so in that essence, it’s a very human-oriented thing still.

Trevor Nicks: But I’ll say if there is an underappreciated part, it might just be how hard scaling something is in terms of the technology and also the organization. It’s very different to when Caravel was just first a four-person company, it really, it kinda feel, still felt like an academic lab that just was talking to companies. But now that we’re ten people and growing, it’s a very different feeling than at least the academic lab that I was in. I know there are larger academic labs than Caravel is, but we’re operating very differently in terms of the flexibility at which ten people are operating together as a unit. It’s not that, oh, this person has one project and they only do that. There is a lot of interplay between everybody’s day-to-day and a lot of interdependence that wasn’t necessarily there when everybody’s doing their own PhD thesis research.

Grant Belgard: What’s been the hardest part of the journey?

Trevor Nicks: For the most part, it’s been pretty enjoyable. There have been hard parts. I think the hardest part has been sometimes simply the uncertainty. Especially, when I started I had a co-founder. Her name is Emily. And we’re still very good friends. But in the end, we pivoted away from some of the technologies that we first thought we were going to develop. And our first contracts were more in the industrial space as opposed to bio health, where we originally thought we might be working. And so she ended up deciding to leave the company and is now, being very successful in a role elsewhere. In that transition period where I went from like having someone who was in it equally with me to being like, “Oh, I’m the one who is in charge of literally everything,” I found out very quickly, oh, I don’t like this where I’m in charge of everything.

Trevor Nicks: And so my goal immediately was to, okay, I have to build another team a support structure that can actually make this not a, oh, this is Trevor and people, but actually this is Caravel. And after Emily originally left, I really spent like the next eight to 12 months finding the people that I felt comfortable bringing into the fold of the leadership in the company and filling out the C-suite. And I’m very grateful for the people that have joined me on that journey because it would be impossible without them. And also the rest of the team too. Like people here are taking on so much responsibility, and I really appreciate it. So if it was just me, it would, one, be lonely, and two, impossible. And so it was really just thinking about curating the organization. And that was extremely difficult. Like there were so many hard decisions to be made during that process.

Grant Belgard: What should computational biologists learn if they want to contribute meaningfully to experimental protein engineering?

Trevor Nicks: That’s a great question. I think you should read more crazy microbiology papers. Because if you’re doing computational biology and you’re taking mostly CS classes and maybe, done some intro-intro to biology, and you’ve learned about, how DNA replication works and how a ribosome works, and you learned that in the maybe the molecular biology class you took. That’s just one model of what exists, right? That’s usually just like what you learn about what’s in the E. coli system and what’s in the human system. Microbes are insane. There are so many different versions of every single thing and there are versions of things we don’t even yet comprehend, and we learn more every day.

Trevor Nicks: And I genuinely think that, some of the statements that some computational biologists sometimes say, it’s not that people are saying things and they’re wrong, but sometimes it feels a little naive in terms of how much compute can do when biology is so weird, and also the idea that we shouldn’t explore more of biology. Sometimes people are saying that. And so I would just say, read, read a paper about archaeal symbionts that live in ocean vents and how they have different membrane proteins than everything else. Read a paper about salt-tolerant organisms that live in the Dead Sea, and these microbes that live in mats at the bottom of the Arctic, and this really cool biology. And then use it as a tool so that you’re not just living in your computer, but you are drawing from all the amazing things nature has already built to build whatever it is that you need to build.

Trevor Nicks: That would be my advice. That’s what I would say would be a useful thing to learn.

Grant Belgard: What piece of conventional career advice do you disagree with?

Trevor Nicks: I think there is I don’t know how conventional this is. Maybe Grant, you can tell me if you think this is conventional. Do you think it’s conventional that a lot of people feel the need to go do their PhDs right after they finish undergrad, especially if they’re like participating in the New England University race?

Grant Belgard: Yeah, that’s pretty conventional.

Trevor Nicks: Yeah, I think it’s– I did that, right? And, Yes, I like worked for, this summer doing s- the algae company Spira. But then I finished undergrad and four months later I was, sitting back and down into a classroom in Boston to do my PhD. And I don’t regret it. But I will say I’ve met a few people and did my PhD with some people. There were people in my lab who were either doing like part-time PhDs while working at, for example, Manus Bio in Boston, and then also working in our lab at Tufts. And I’ve met people who, were in industry for six to ten years before then going and doing their PhD. And those individuals brought such a perspective that I didn’t have and such a depth of knowledge and understanding to not only how the world works, but all like these systems they had been studying in and the way in like this hierarch, like the hierarchy of systems that I didn’t really yet appreciate.

Trevor Nicks: And so I would just say you don’t have to immediately go to grad school if you want to go to grad school. It can be really useful, I think, to go learn something elsewhere. And within that too is you might learn that like maybe you don’t actually want to go all the way through with a PhD. Maybe a master’s would get you where you want to go. So that’s something I would think about.

Grant Belgard: Trevor, what are you Most excited about in the next 12 months?

Trevor Nicks: I am really excited to tell people more about what we’re doing at Caravel. In the first three years of our company, it was very yeah, we had a website and we were talking to the NSF and to customers and stuff, but we weren’t being very public about it. One, ’cause it was so early. Two, we had to file IP on these new processes that we were working with. But now we’re at a moment in time where we have this new technology. It works well enough that we’re like we think we can build some really cool things here. Let’s talk about it more.” And then at the same time too it’s a hopeful thing in that biology’s been in a little bit of a rut in terms of our ability as a field to deliver really scalable products and at the same time financial returns.

Trevor Nicks: And so we’re a little bit hopeful that this new system will allow us to, again, generate more data and generate data for less cost in a way that will allow more people to build scalable bioproducts. And I just wanna have a little bit of hope, a little bit of joy maybe in the conversation around the future of biology as a field and our ability to biologize more of industry. So that’s what I’m excited about.

Grant Belgard: Where should listeners go to follow your work or learn more?

Trevor Nicks: High level, our website is Caravel, C-A-R-A-V-E-L.B-I-O. We’re named after the Spanish and Portuguese ship that helped circumnavigate the world for the first time. And that idea of exploration is baked into our DNA. My husband is also Spanish. Beyond that, LinkedIn, of course, is a place where we share a lot and perhaps we’ll set up a blog soon. We’re also on Instagram and Twitter, now X @caravelbio

Grant Belgard: Trevor thank you so much for joining us.

Trevor Nicks: Yeah. Thank you, Grant. Appreciate it.

Shannan Ho Sui episode coverart

The Bioinformatics CRO Podcast

Episode 92 with Shannan Ho Sui

Shannan Ho Sui, Principal Research Scientist in the Department of Biostatistics at the Harvard T.H. Chan School of Public Health and Director of the Harvard Chan Bioinformatics Core, discusses what bioinformatics core facilities do and how new technologies let us answer new research questions.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Shannan Ho Sui is a Principal Research Scientist in the Department of Biostatistics at the Harvard T.H. Chan School of Public Health and Director of the Harvard Chan Bioinformatics Core.

Transcript of Episode 92: Shannan Ho Sui

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. Today, I’m speaking with Shannan Ho Sui, principal research scientist in the Department of Biostatistics at the Harvard T.H. Chan School of Public Health and director of the Harvard Chan Bioinformatics Core. The core supports researchers through bioinformatics analysis, training, and platform development with a focus on high-throughput sequencing and reproducible collaborative research.

Grant Belgard: Shannan’s path spans biochemistry, genetics, iPSCs, pathogen genomics, cancer genomics, and now the leadership of a team working across many areas of modern computational biology. Today, we’ll talk about what a bioinformatics core does, how she built her career, and what advice she has for scientists working at the interface of biology, computation, and collaboration.

Grant Belgard: Shannan, welcome to the show.

Shannan Ho Sui: Thank you so much, Grant. It’s a pleasure to be here

Grant Belgard: So for listeners who’ve never worked with a bioinformatics core, how would you describe your current role?

Shannan Ho Sui: As the director of a bioinformatics core, we’re really there to support researchers with their data analysis. And so my role is to sit at that intersection between biology, statistics, and computing, and find solutions to their biological questions that they have from these large, high throughput, messy data sets. being at the core means that I have to put on a number of different hats at different times. As you already mentioned, we focus on consulting which involves analysis itself training, and also developing platforms. And so not only do I have to wear the different hats in terms of developing curricula, but also developing analysis strategies and also trying to think about reproducible ways to run pipelines.

Shannan Ho Sui: But there’s also that component of interacting with people and trying to really listen and understand what their biological question is, so that we can find the right methods that are appropriate for their data, while also managing sort of the finances of the core resource allocation. So a lot of different things that I have to do in my role.

Grant Belgard: What kinds of scientific conversations are you most excited to have right now?

Shannan Ho Sui: It’s an interesting question. I think there’s a number of different themes that have come up recently that have been very interesting. I’m, by training and at my heart, I’m a biologist, so the conversations I end up having are the ones that have a really interesting biological puzzle to solve, while at the same time, new technologies are emerging, and especially in the spatial and single-cell spheres.

Shannan Ho Sui: In the last few years, we’ve been doing a lot of work in that area, and it’s allowing us to answer these very interesting biological questions in new ways. And we have some great collaborations within the Harvard community. Scientists are doing really innovative, interesting work. Some of the work we’ve been doing recently is with Dr.

Shannan Ho Sui: Rachael Clark at Brigham and Women’s Hospital, who’s a dermatologist, but does a lot of research into inflammation and immune responses in skin. And so I’ve worked with her for, I think it’s 10 years now. But we’re still doing really interesting things looking at face transplant rejection and now being able to use spatial technologies and single cells to really characterize those pro-inflammatory and anti-inflammatory responses when you have an a transplant, and it translates to other transplant tissues as well. And also, I enjoy having conversations about given a specific biological question you’re trying to answer, what is the best technology to try to tackle this, and what is feasible and what is not? Because I think a lot of times researchers have great ideas, but it’s not always clear what the approach is to try to get at those questions.

Shannan Ho Sui: And sometimes what they think they want is not necessarily, the most appropriate. Trying to solve those puzzles is what I find most interesting.

Grant Belgard: What’s a misconception people often have about bioinformatics support?

Shannan Ho Sui: I think one of the things that we really strive to do in our core is to be at the level of a collaborator. So even though sometimes, you know, a bioinformatics core is classified as service, which it is also a collaboration. Very rarely I’ll be able to just run things or set an analysis pipeline and hand back the results.

Shannan Ho Sui: It’s very interactive. Um, we bring a lot of expertise both on how do we analyze that data, but also how to interpret it. so I think, I think that’s maybe a misconception. I think there are still folks out there who think, “Oh, I have a data set. I’ll just send it off. I’ll get some answers, and I’ll be off running, writing my paper.” especially with the complexity now that we’re getting from these new technologies. We plan to talk to our researchers a lot about what we’re seeing, how to interpret, we come, too, with a lot of biological knowledge. Within my group everybody has a biology background. And so there’s a real component of interpretation that’s key to supporting projects

Grant Belgard: What makes a project especially well-suited to a bioinformatics core?

Shannan Ho Sui: I think the best projects have a very clear biological question or hypothesis, even if we’re not completely sure how we’re going to analyze that data. It’s clear what we can define as success. The other thing that makes it well-suited for the core is if it’s something that we are already familiar with and something that we see frequently and have methods in place that are best practice. That’s not always the case, and sometimes we do projects where new technology is coming onto the scene and would benefit from the expertise that we can transfer over from other experiences we’ve had. but in those cases, it’s only suited for the Bioinformatics Core if the collaborator understands that there’s a component that is learning and development in addition to the analysis itself. And so I think that all translates into what I was saying before about it being a true collaboration.

Shannan Ho Sui: Need to be able to sit down, discuss the design, to iterate on results, and to invest time in communication, not just send data and get a figure

Grant Belgard: Related to that, what does an ideal first conversation with a collaborator sound like?

Shannan Ho Sui: I always start the conversations that I have with our collaborators with asking them what is your biological question?” Usually they come to me, and they already have a study in mind that they either are planning or are have already generated data for. And for me, the background on the biology is key to the whole thing being a success. So at the outset, I would like to know, what is the motivation for this study? How did you get to where you are now? And, what preliminary results or previous results led to you asking this question in the first place? And so then from there, I wanna be able to talk about… ideally, they would have not done the experiment yet, and they would be coming to me for a design consultation. And so then I will ask them, “Have you already thought about how you’re planning to answer this question?” that usually gets us into weeds of, what is the source material?

Shannan Ho Sui: How many replicates are they doing? What are the conditions? Do they have the appropriate controls in place? What are the key comparisons that they’re planning, to execute? And I think in that conversation also, just being frank about what we can do and what we can’t do, and where there are risks and where there are things that we think will be fairly straightforward. being able to talk about those things in a very open and candid way is important, and I think during that first conversation, if it’s an ideal conversation there will be a shared understanding that this is research, that bioinformatics is not going to solve all the problems.

Shannan Ho Sui: It’s going to help answer questions but we’re going to have to work together to get that done.

Grant Belgard: What are the earliest signs that a project could run into trouble?

Shannan Ho Sui: Oh, you know, just talking about this with my team recently because I think we’ve been re-scoping a couple projects, that went over budget, and one of the red flags I think is when… first I’ll say that estimating how long a project is going to take is incredibly difficult without knowing or having seen the data or the QC on the data.

Shannan Ho Sui: So I always tell people, “I don’t know exactly how long this is going to take. This is a difficult thing to do. We’re going to have to be willing to accept ranges, and we’re also going to have to accept that there are going to be these natural stopping points to check in, and we may need to reassess budget at that point, depending on what we’re seeing in the data.” If I get pushback on that or if it seems like the collaborator’s potentially inflexible on budget or inflexible on, know, how the analysis is going to be done that usually tells me that we’re likely to run into problems.

Grant Belgard: What you wish researchers would ask before generating data?

Shannan Ho Sui: I wish they would come to me to talk about the experimental design. Just having that detailed discussion on what they’re planning to do. I also sometimes … I’ve had people ask me this, “What do I need to do, as an experimentalist to make this collaboration go well?”

Shannan Ho Sui: Because there are things you can do, like being very well organized with your metadata is a key component of having a successful collaboration with the Bioinformatics Core. And so asking what format we need it in, that, that’s helpful. I think also, I always ask about timeline for the project because I want to know if there are upcoming deadlines. it would … It’s always helpful if they ask me too what is the turnaround time? And also h- let me know, this is not a question, but informing me if they expect to have long breaks in between. Because as you can imagine in academia, sometimes you do an experiment, you get it analyzed through the core, but then you may be preoccupied with other things that have come up and you’re juggling multiple priorities. analysts, it’s much more efficient for them if there’s a quick turnaround from their side, too.

Shannan Ho Sui: So having that discussion about timelines and what we can anticipate in terms of the communication and turnaround times is super helpful.

Grant Belgard: When, uh, they’re not submitting the paper three years later after lots of follow-up and coming back to you

Shannan Ho Sui: Yeah, no, that’s true. But also just, and I fully understand this. We have a lot of clinician scientists who we work with. When they tell me they’re going to be going on clinic for five weeks, I know I’m probably not hearing from them for that period of time, and that’s fine. just so long as we know that, okay, we’re, as a core we work fee for service, and it’s paid hourly.

Shannan Ho Sui: So interim period, I’ll be trying to find some small project or something well-defined that person can work on in the meantime

Grant Belgard: What kinds of questions are hardest to translate into an analysis plan?

Shannan Ho Sui: I think some of the more exploratory questions can be difficult to translate. Yeah so there are some things that are standard and best practice pipelines for analyzing data. So I’ll give an example of a spatial transcriptomics project that we worked on recently. it was a large study, I think something 20 to 30 slides. And it did have a hypothesis, and I’m not gonna I’m not gonna say exactly what the project was because I wanna protect our collaborators. But it was done more just to see what they might see out of the project rather than I have a clear hypothesis about a particular mechanism or a pathway that I’m interested in interrogating.

Shannan Ho Sui: And so while you can do all the standard things of QC-ing the data quantifying, doing differential expression, looking across time points or looking across slides Without knowing what they really want to focus on, and this was a brain study, so without knowing whether, it’s a particular region or a particular type of cell that they’re interested in, you’re going all over the map just looking for anything that the data can show you.

Shannan Ho Sui: And sometimes you’re lucky and there’s something very clear and interesting that pops out. But there are many times where there’s lots of different things that could be interesting or sometimes very little effect size, for example, in a study with a treatment. And so there it becomes a question of, okay now that we’ve gotten through to this point and we ha- we need to decide, what the hypothesis is and what we’re gonna focus on if there isn’t a clear objective, then the analysis plan becomes very difficult to continue with, so

Grant Belgard: What does successful collaboration look like from your side?

Shannan Ho Sui: One of the things that I’ve really valued in some of my collaborations is, number one, trust and mutual respect on both sides. Recognize that our collaborators come with, come to us with tons of amazing expertise, and they’re being incredibly innovative with their datasets. and if they’re able to be organized, like I was saying about the metadata and the data itself, then trust us to take that data and t- look at the quality and try to preserve as much of it as we can that is reasonable because as I mentioned, we worked a lot with clinicians.

Shannan Ho Sui: So sometimes, you don’t have a perfect experimental design, or you don’t have perfect samples. You’re working with what you have, and they are precious samples. I think if there’s trust that we’re going to try to get the most signal we can out of the data while also being rigorous and being objective about what really is, high quality enough to move forward with, I think that really helps the collaboration. And then the part that I’ve maybe value the most, which is then when we provide the results back, that there is a lot of discussion and interpretation on both sides, and that back and forth between this is, from the bioinformatician, this is what I’m observing. This is what I think is true signal, and then that being interpreted in its context.

Shannan Ho Sui: So for example, being like, “Oh, that’s a T cell signal, and I was expecting that because of X, Y, and Z.” Those, that I think if you can get the bioinformatician to understand the biology that they’re discovering through the collaboration and build domain expertise over time, which just makes them better and better, while at the same time providing the collaborator back with scientific insight into their dataset, I think that’s what I would consider a success.

Grant Belgard: What parts of the work are most visible to collaborators and what parts are most invisible?

Shannan Ho Sui: The most visible are all the visualizations and the figures, the nice pictures that we create to try to help them understand and also then for their papers. most invisible is all the challenges that you have when working with a bioinformatics dataset; often there’s a lot of data cleaning that has to happen. Sometimes I think it’s really obscured how difficult it can be to run an algorithm how long it takes to run a particular algorithm. I was just doing an InferCNV analysis on a single-cell dataset recently, it’s a large single-cell dataset. Tweaking those parameters to make sure that it’s the result is rigorous and reliable, each run was taking two to three hours.

Shannan Ho Sui: and sometimes it was running out of memory, and those are not things that, you necessarily share with your collaborator that, oh, I gave it 120 gigs, and that still wasn’t enough, and then I had to re-run again for three hours. So I think that component is not as visible, and I think the component that sometimes I wish there was a way to resolve this is that, within my team, we have a lot of internal discussions about the datasets we’re looking at.

Shannan Ho Sui: Yes, there is a primary analyst who’s been assigned to the dataset, but there’s almost always a few people involved in helping to look at that data and making sure that the interpretation makes sense or if there’s been a, an issue or a particularly sort of decision point that needed to be made other bioinformaticians weighed in on that.

Shannan Ho Sui: And I think by the time we send the report, it looks like everything was straightforward and easy, and it looks like it has these pretty pictures in it that make a lot of sense. But it probably took much longer than the collaborator thinks it did to get there.

Grant Belgard: What do you think about authorship credit and intellectual contribution in core supported work?

Shannan Ho Sui: I’ve been the director of this core now since 2015, and when I took it over, I think that co-authorship wasn’t as emphasized as it has become over the last 10 years or so. So when I became the director, I made a conscious effort to emphasize that, at the level that we’re collaborating and the amount of intellectual input we’re providing, that we are going to request co-authorship. And so over time I would say that for a majority of papers that get published where we’ve worked with collaborators, we are getting co-authorship. Obviously, that’s somewhere in the middle of the authorship list because we’re not the data generators. But I think it has become better understood that, without the bioinformatics expertise, a lot of these studies wouldn’t be published.

Shannan Ho Sui: And so it is a, a critical component of the collaboration. I’m not shy to ask anymore.

Grant Belgard: Yeah. What’s one thing listeners might be surprised to learn about running a bioinformatics core?

Shannan Ho Sui: I run my core like a small business. It’s like a startup. There are … All of our income or revenue is generated through cost recovery. So the amount of time we spend on a project is billed in hours, and the major resource or the major expense are people. And so being able to balance the books and to, to make that work, because as a NIH approved core, we have to break even.

Shannan Ho Sui: We’re a nonprofit, essentially. So we have to be within that 15% of break even to meet requirements. And so a lot of financial management. There’s a lot of project management, resource allocation. We do get involved in writing grants. And then there’s, that component of um, outreach and marketing and business development that I think that, you wouldn’t necessarily think about when thinking of a core facility.

Grant Belgard: What does reproducibility mean in day-to-day bioinformatics work?

Shannan Ho Sui: I think this is something that is really difficult and something that we’re always striving to improve upon. For me, reproducibility means that if you do a project and you’ve executed an analysis, that if somebody comes back several years later, that you can, one, you know where that project is and you can You still have copies of that data, but then you can run the pipeline that you ran on that data and get a very similar result. I’m not gonna say the exact same result, because tools change, versions change, and things like that. But in essence, that you can replicate or reproduce that analysis. for us, that means, good documentation in the code, committing the code and versioning it in GitHub, writing our code in things like Python notebooks or R Markdown, along with the interpretation so that it’s clear which figure was used to, come to a particular conclusion.

Shannan Ho Sui: yeah, having the data be safe somewhere and accessible, as I said before. We’re always working on this. I think it’s something that is overlooked a lot the time, and I also think it’s something that doesn’t come naturally to people, and you have to have processes and standard operating procedures in place for it.

Grant Belgard: Where do reproducibility problems usually enter an analysis?

Shannan Ho Sui: I think the long-running complex projects are the ones that are most difficult to reproduce especially when you have a rich data set, you might pursue many different

Shannan Ho Sui: analyses

Shannan Ho Sui: to answer a particular question and iterate on that, changing the parameters and things slightly, at which point it becomes really tricky to know, what you need to save and what you need to eliminate. And Making it clear what path was taken to get to the final result. But even in smaller project, I think reproducibility can be challenging. One of the things that the core struggles with is how long to keep data in a project. Theoretically, we’re not responsible for that data, the person who generated it is. And as you can imagine, as a core, we have a lot of people’s data, so we’re using up a lot of storage space. And so we do have to make decisions about when we let that data go, when we finally delete it.

Shannan Ho Sui: Deciding when a project is over a, is not straightforward at all because people can come back years later asking for a reanalysis or even just the data so that they can finally submit their paper. And so I think challenges can come really at any step there, at any step of, when you delete the data, where you put the code, which version of the code you, you keep as the, the final version. But even things like, for example, we use templates a lot in our group so that we can standardize on what we think is best practice for certain analyses. Sometimes there can be a problem just because somebody didn’t modify something sufficiently in the template to match the data because there was a copy and paste error or something. I think it’s just a really challenging thing, but something that we definitely strive to try to address in my group

Grant Belgard: How do you handle the tension between rapidly changing methods and the need for stable, reproducible results?

Shannan Ho Sui: That’s a tricky one. As a core we need to be able to standardize and but at the same time we also need to be able to keep up with emerging methods. So we have a lot of discussions internally about what best practice is, but we also work a lot with the community. And we are part of, for example, the NF Core community.

Shannan Ho Sui: We have been part of Bioconductor. We wanna know what other people are doing and to maintain what is the current state-of-the-art in analysis. But every time you change your pipeline or change your approach, it’s extra time and extra effort, not just to try a new tool or to benchmark it for an existing dataset, but also to then make that reproducible for the future. For the most part we are using, as I mentioned, our templates and what we have established as best practice, but we’re always on the lookout for what looks like it might be emerging as the next phase or the next best practice for an analysis. We’re fortunate to be in a community where

Shannan Ho Sui: there’s

Shannan Ho Sui: a lot of communication, there’s a lot of seminars, there’s a lot of discussion about what people could be or should be doing with datasets.

Shannan Ho Sui: And so I think you can never be complacent. You can never think that you’ve got it solved, but you do have to standardize and, … So I would say 80% is probably reproducible pipelines, and then 20% is pushing forward all the time to try to incorporate new things.

Grant Belgard: What should wet lab biologists understand before starting an omics experiment?

Shannan Ho Sui: I think one of the key things is before you do the omics experiment, is that what you need to answer your question? Some things could be solved with a qPCR or, with something much less expensive. And I think just there are, depending on the types of omics, being aware of things like batch effects and confounding and, Also being aware of, you can generate these really exciting datasets, but you also need to budget appropriate time analyze those datasets because I think a lot of times it’s very exciting to try a new technology. but if it is truly an emerging technology and has just come out on the market, there are going to be very few, standardized pipelines to analyze that data. And so then just being aware that they’re going to need to budget extra time so that somebody can really dig in and learn how to analyze that dataset well.

Grant Belgard: What bioinformatics concepts tend to unlock the most value for experimentalists?

Shannan Ho Sui: So the whole design, I think. So definitely understanding variation. When I do my consults, I will talk about replicates, but I’ll talk about that in the context of the variation that we might be expecting. So if they’re doing a clinical study, I know that there’s going to be a ton of biological variation.

Shannan Ho Sui: You’re going to need high numbers of replicates something from a cell line, for example. I think variation probably is of the highest value things that they can understand because even when you do time course for analysis there are different ways to set up a time course. You don’t always have to start at T zero for everything, right?

Shannan Ho Sui: You could end everything at the same time point, for example, and that might make sense in some cases. Yeah, I would say like technical variation and batch confounding, are high value because they can really make or break an experiment.

Grant Belgard: So we’re about to switch course for a bit and talk about trends in science and technology. But I just wanted to comment that so far your answers, if the tables were turned and you’re asking me the same questions you know, I would give very similar answers, probably not as eloquent, but essentially the same content, right?

Grant Belgard: I think everyone involved in providing bioinformatics as a service for long enough runs into the same issues time and time again.

Shannan Ho Sui: Yeah, no, it’s true. And I think yeah, I think if you’ve been in this field for a while, you would’ve seen how those decisions about the design really impact your ability to interpret that data. And yeah I would imagine that pretty much all core directors would be saying things very similar to me.

Grant Belgard: Which kinds of biological questions feel newly approachable because of current omics technologies?

Shannan Ho Sui: Yeah, so I think spatial is really giving us a lot of insight into things like cell-cell communication and local niches. And I think that’s really exciting. I think, single-cell was exciting because we had this ability then to get higher resolution into what might be happening, and we could look at signaling between cells, but it was still noisy because you couldn’t be sure that the cells were really in close proximity to each other or even in contact with each other. And so now I think that’s giving us the ability to ask some very well-defined questions. Again, I come back to this idea of if you have a really well-defined question, then that’s gonna help make the project a success. And so for example, if you have a hypothesis about, particular immune aggregates interacting with, maybe parasites or endothelial cells, you can now look at that.

Shannan Ho Sui: You can… and you can, subset your dataset to a particular niche and look very carefully at that. Whereas before we would, we could look at it at single-cell resolution, but it was all mixed in together. So I think we’re gonna make some very interesting findings using those approaches, and I’m excited about that.

Grant Belgard: How can scientists decide whether a more complex assay is actually worth doing?

Shannan Ho Sui: Yeah, so when we meet with people, that’s one of the key things that we’re always aware of. A lot of these studies are very expensive. so there’s always this balance of what is the question? What is the most straightforward way of getting to an answer for, to that question, and is it worth the cost?

Shannan Ho Sui: What are the potential pitfalls of pursuing a particular technology? Sometimes there are things that can create issues, like if you’re looking at a rare cell type and you’re taking a tissue section, how… i’ll ask very detailed questions like what proportion of those cells in the slide do you think you’re going to be able to detect? And how variable is that from slide to slide? Is that going to be in every slide, or y- do you just have to get lucky and get it just right to find what you’re looking for?” think, There are older techniques that not being used anymore that maybe could still be helpful instead of doing a very expensive spatial study. I haven’t had anybody doing laser capture dissection lately. I haven’t had any of those, but that was quite powerful for a while.

Grant Belgard: It is, yeah

Shannan Ho Sui: yeah. So I think it’s really thinking about the question and then finding the method rather than getting excited about a, a technology and… there is, as a bioinformatician, I’m always excited when somebody brings me something new to work on, but Weighing the cost of that with the, the return on that investment is important.

Grant Belgard: How can researchers make their data sets easier to interpret, reuse, and build on?

Shannan Ho Sui: Yes, I already mentioned the metadata. I think metadata is key. and lot of transparency. So I, there are the, the regular things that you can put in the datas, in the metadata, the types, the sex, the obviously. But then there’s all kinds of other information that the experimentalists can share.

Shannan Ho Sui: For example, the date of extraction for each sample, when were all the libraries prepared, even if it was done by two different people can have an impact. As much information as people are willing to provide, I’ll take all of it. So I think that is the key to the interpretation as well, because, the data will give you clues and tell you.

Shannan Ho Sui: We’ve had instances where we look at a study and we see something unusual in a couple samples, and when we go back to the researcher and we start to dig in and ask more questions, that’s when we find out, oh, for example, those two were the last ones that went onto the machine, right? Or onto the instrument.

Shannan Ho Sui: And so those, that sort of information upfront can help us so that we’re not spending time trying to figure out what the problem was, but we already know, okay, so they suspect that potentially these two samples might be different because they know they were last and there was like a, a bit of a longer break than they wanted between the first and the last samples.

Shannan Ho Sui: All of that aids interpretability. I think for single-cell data, domain expertise that people bring to the table are incredibly helpful. It’s getting easier to cell type with more datasets available in the public domain and also with

Shannan Ho Sui: AI.

Shannan Ho Sui: it still doesn’t beat domain expertise that biologists have built up over years. So telling us, we’re expecting these particular cell types, these are the markers that we trust most reliably for this, that helps a great deal.

Grant Belgard: How did you first find your way towards bioinformatics?

Shannan Ho Sui: I started in undergrad, I did my degree in biochemistry and molecular biology. And in my honors thesis rotation, was sequencing archaebacteria and looking for sites of RNA methylation. And I’m gonna date myself now, but, back then, a sort of bioinformatics involved running BLAST searches and sequence alignments.

Shannan Ho Sui: And I just found it incredibly interesting that you could identify what species you were working with from a couple of these searches using your computer, and it was very instantly gratifying. And around that time, I had a lot of friends who were in computer science. And so I was seeing them program, and I had started becoming interested in programming as well. And so then at that time, I actually thought I was gonna go to med school, and I was preparing sort of pre-med but decided at the last moment to take a second degree in computer science. And at that point I already knew that biology was what excited me and that I wanted to contribute to research. But it very quickly became apparent to me that I enjoyed programming and I enjoyed being able to answer some questions quickly, and that I wasn’t necessarily the best at lab work and going in on the weekend to check on my cells.

Shannan Ho Sui: and so I ended up moving into computational biology. So what had happened at that time is that my timing was fantastic because as I was wrapping up my second degree, that’s when–

Shannan Ho Sui: I’m from Canada, so I was at Simon Fraser University. But Simon Fraser and the University of British Columbia were putting together their, the very first cohort for a bioinformatics PhD program. And I happened to be designing a database for Fiona Brinkman, who was one of the people the PhD program, and she said, “You should definitely apply for this and move towards PhD in bioinformatics.”

Shannan Ho Sui: And so that’s what I did. And it was, know, best thing I ever did. I have so enjoyed it and have over that, this entire period excited about the work and the breadth of things that I can do, the variety.

Grant Belgard: Subsequently what were the key turning points in your career?

Shannan Ho Sui: Yeah. I did my PhD in genetics and bioinformatics. At that point, it was not even clear that bioinformatics would be a real discipline that you could do a, a PhD in. And my degree is actually in genetics, but it was all computational. I then had to decide what I wanted to do after the PhD, which was focused on gene regulatory networks and I ended up taking a position that involved a combination of research, but also project management.

Shannan Ho Sui: So I ended up being a, a scientist and a project manager for an initiative called Bioinformatics for Combating Infectious Diseases. And so that was my first foray into working in large collaborative teams and managing that. And and so that involved 11 different researchers in that consortium, and then having to understand what everybody was doing and corral people towards a common goal, was a great experience and it made it… know, it was with Fiona Brinkman. I actually went back to her after that to work in that position. It really showed me how, as a leader or a manager, you can leverage different e-expertise to do more than you could ever do on your own. And so after that, I ended up transitioning into Boston in working with Winston Hide in stem cell research.

Shannan Ho Sui: And so I was going from one domain to another domain, seeing all these different transferable skills, and then I had to make the decision of whether I wanted to, I think, have my own lab and pursue like a purely academic career or the opportunity came about that I could direct a core facility. And by that point, I’d already seen cores had access to such a wide variety of data and such a, an interesting set of questions that people were answering. And to be frank, I never had one particular interest in research that I wanted to pursue. And so I probably would’ve been a really bad professor because I didn’t have my own thing that I really was excited about.

Shannan Ho Sui: I was excited about what everybody else was doing. so that became the turning point where I essentially made the decision to direct the core and to support other people’s research rather than focusing on my own interests, which as I mentioned, I didn’t have a clear idea of what that was anyway.

Grant Belgard: Pivoting to issues in running a core, uh, what makes someone excellent in a collaborative bioinformatics role? What do you, what do you look for when you’re hiring?

Shannan Ho Sui: It’s a combination of a lot of different things. And I think it’s not just for cores, it’s for every and but it’s especially important in cores. So soft skills, I think, are incredibly important when you’re working in a core because you have to be able to really listen and understand what what people need for their projects.

Shannan Ho Sui: And so that involves, soft skills, deep biological knowledge, and technical ability. Those three things are incredibly important. I don’t think you can really be in a core without all three of those. And so I think maybe that’s also just the way that I’ve run the core that I want people to have that diversity in their repertoire.

Shannan Ho Sui: I know probably there are other models where you delegate specific things to people where their strengths are. For example, one person being more technical and then another person handling the communication. But for me, I think to be able to do all three of those things is incredibly important and valuable, and it just makes things much smoother.

Shannan Ho Sui: One of the things that I’ve also seen people excel is the ability to teach, not just communicate, but teach. And we have somebody on our team who is so articulate and eloquent at explaining things that it just makes the collaborations run very smoothly. And and I’d also say probably when it- in the realm of soft skills, the ability to take criticism and to also be able to stand your ground when you know that your approach is probably going to be the better one while at the same time being flexible.

Shannan Ho Sui: So I sound like I’m saying a lot of contradictory things, I think. But both flexibility and the ability to hold your ground, I think, are important if you’re going to collaborate.

Grant Belgard: What leadership lessons from bioinformatics cores would transfer well to biotech, pharma, or academic labs?

Shannan Ho Sui: Through my career, one of the things that I’ve found that people have found appealing from the experience I’ve gained as a core director the ability to build teams. So to understand what the needs are of the team and then to recruit people to fit those different roles. And so I’d say that that skill is obviously transferable to biotech to, to be able to resource small teams effectively so that they can be v- be very efficient and effective and successful.

Grant Belgard: Now the million-dollar question AI. How has it been impacting what you do? What are your expectations for where all this will go? What do you expect?

Shannan Ho Sui: Yeah. You and I had a brief conversation about that earlier. I think AI has a lot of potential benefits for bioinformatics. I think it’s great at helping people be better coders or to implement and execute code more quickly. Helps with the reproducibility, and there’s a lot of places where it can help with, you know, single-cell, for example, cell typing or anything that’s repetitive.

Shannan Ho Sui: I think AI or anything where you’re summarizing information, AI is incredibly helpful and a huge resource. It’s not particularly creative. So I think that in places where you’re looking for something a little bit more creative or more innovative AI is not going to replace that, although it can help make the more busy work or less interesting things faster. And as I mentioned in our previous conversation, one of the things that has come up is that while it’s speeding up certain components, it’s actually slowing us down in other ways. And a lot of our best practice for analysis has been put together through years of experience and through the community and talking to people and looking through t- at lots of different data sets with different nuances.

Shannan Ho Sui: But we’ve had occasions where collaborators take, the approach or the code, put it through Claude or ChatGPT, and then feel a sense of mistrust in the analysis because AI is suggesting something different. And I think that is something that we’re gonna have to learn to deal with and how to cope with that because, a lot of times AI will generate all possible solutions for a problem. And the one that we’ve selected is one that we likely have felt confident about, but now we have to explain why we didn’t choose all of these other ones. And that takes more time than I think that we often have and since we’re charging by the hour, it’s not really a good use of funds.

Shannan Ho Sui: So From that perspective, I think again, this is a communication and a p- and a people challenge, not necessarily an AI challenge. So teaching people how to use AI responsibly, which is something that we’ve been brainstorming coursework for recently. And we’re– we’ve already had a couple of courses on this, but we wanna flesh out that program a bit. But on the other hand, I think we’re going to have to use it. Like we… and I think that it has a lot of potential to make, our work easier and also to us better as bioinformaticians. So just really understanding it, understanding its limitations, understanding what it’s good for and where might want to be more critical is gonna be important.

Shannan Ho Sui: I often think, ’cause I, I have a young child what are the skills that we need to be teaching our young children for the future if AI is going to make many of the things we do now, easier and more automated? And I keep coming back to critical thinking. And for bioinformaticians, I think that means having deep biological knowledge because I don’t think you can think critically about biological problem without that.

Grant Belgard: Great answer. Shannan, thank you so much for joining us.

Shannan Ho Sui: Yeah, thank you so much for having me. It’s been a pleasure, Grant.

The Bioinformatics CRO Newsletter

The Bioinformatics CRO Newsletter

August 2026

Data to Discovery: What’s New at the BCRO

Welcome to the first edition of The Bioinformatics CRO Newsletter.

Every quarter, we’ll share recent conversations from The Bioinformatics CRO Podcast, selected publications from our bioinformaticians and collaborators, practical guidance for working with complex biological data, and examples of how bioinformatics is helping research teams move projects forward.

In this edition: clinical genomics, precision immunology, recent work across proteomics and microbial genomics, an example of bioinformatics in practice, and a practical guide to preparing project metadata.

New on The Bioinformatics CRO Podcast

Elliott Margulies: Molecular diagnostics and clinical genomics

Elliott Margulies, Director of Bioinformatics at BillionToOne, discusses the application of molecular diagnostics to clinical testing and a career spanning genome technology, computation, and clinical translation.

Listen to Episode 87

Elliott Margulies - The Bioinformatics CRO Podcast

Jason Mellad - The Bioinformatics CRO Podcast

Jason Mellad: Building platforms for precision immunology

Jason Mellad, co-founder and CEO of OtoImmune, shares his vision for personalized precision immunology and explains how patient need can shape new approaches to immune and autoimmune disease.

Listen to Episode 85 →

Dr. Eric Green: The Human Genome Project and the future of genomics

Dr. Eric Green, Chief Medical Officer at Illumina and former Director of the National Human Genome Research Institute, reflects on the Human Genome Project, his time leading NHGRI, and the future of clinical genomics.

Listen to Episode 84 →

Eric Green - The Bioinformatics CRO Podcast

Recent Publications

Our bioinformaticians continue to contribute to research across a broad range of scientific areas.

Integrating deep learning for post-translational modification crosstalk on Hsp90 and drug binding

This 2025 Journal of Biological Chemistry study combined mass spectrometry and deep learning to investigate how patterns of post-translational modification on Hsp90 may influence ATP and inhibitor binding.

Read the publication →

Polybacterial intracellular macromolecules shape single-cell inflammatory profiles in upper airway epithelia

Published in npj Biofilms and Microbiomes, this study examined how bacterial macromolecules within epithelial cells can shape inflammatory signaling across upper-airway mucosal sites.

Read the publication →

The first single-cell sequencing of Plasmodiophora brassicae reveals genetic diversity and clonal dynamics

This Frontiers in Microbiology paper reported the first single-cell whole-genome sequencing of a plant pathogen and identified substantial genetic diversity within cells isolated from a single infected root.

Read the publication →

Practical Bioinformatics Tip

Why metadata is often the most important file in your project.

Sequencing files contain the biological measurements, but the metadata determines whether those measurements can be interpreted correctly.

A well-prepared metadata file connects each sample to the information required for analysis. Depending on the project, that may include treatment group, disease status, collection time, tissue type, experimental batch, sequencing run, patient or animal identifier, and relevant clinical or experimental variables.

Incomplete or inconsistent metadata can delay a project, limit the analyses that can be performed, and make it difficult to distinguish genuine biological signals from technical variation.

Before transferring your data, we recommend:

  • Using one row per sample and one column per variable
  • Ensuring sample names match the sequencing files exactly
  • Including experimental groups, batches, time points, and relevant covariates
  • Defining abbreviations, units, and coded values
  • Clearly identifying missing values
  • Avoiding merged cells, hidden formulas, and colour-based annotations

A clean metadata file can save time, reduce errors, and help your bioinformatician design a more rigorous analysis from the beginning.

Stay in Touch

If we’ve worked together before, thank you. We’re grateful for the opportunity to support your science.

As your priorities evolve, we’d be glad to reconnect. That could mean scoping a new analysis, reviewing an experimental design, extending a past project, or simply talking through what may now be possible with your data

The Bioinformatics CRO Podcast

Episode 91 with Zameel Cader

Professor Zameel Cader, a clinician scientist and Professor of Neuroscience and Neurology at Oxford University, discusses his work at the intersection of headache and pain disorders, human genetics, induced pluripotent stem cell models, blood-brain barrier biology, and translational drug discovery.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Zameel Cader

Zameel Cader is Professor of Neuroscience and Neurology at Oxford University and co-founder of Oxford StemTech.

Transcript of Episode 91: Zameel Cader

Disclaimer: Transcripts are automated and may contain errors.

Grant: Welcome to the Bioinformatics CRO podcast. Today, I’m joined by Professor Zam Cader, a clinician scientist at the University of Oxford. Zam is professor of neuroscience and neurology, a consultant neurologist, and director of the Oxford Headache Center. His work sits at a fascinating intersection: headache and pain disorders, human genetics, induced pluripotent stem cell models, blood-brain barrier biology, and translational drug discovery.

Grant: His group works on building more human relevant systems for understanding neurological disease and for finding better therapeutic targets. What makes Zam’s story especially interesting is the breadth of roles he has taken on: physician, neuroscientist, consortium leader, and entrepreneur. Today, we’ll talk about what he’s working on now, how he got there, and what advice he has for clinicians and scientists who want to build careers that cross disciplinary boundaries.

Grant: Welcome to the show, Zam.

Zameel Cader: Hi, Grant.

Zameel Cader: Great to be here

Grant: For listeners who are meeting you for the first time, how do you describe the scientific and clinical problems you spend most of your time thinking about now?

Zameel Cader: So I think both shape each other a lot. I think one of the fantastic things about being both a clinician and a scientist is that you’ve got a ready motivation to keep you coming in, keep working on the problems that you keep working on. So I’m a neurologist so I see people with brain and nerve problems, one of the commonest things that I see are people with headache and with pain. And it’s such a burden, both on the individual but also on society, on healthcare systems. We do have pain drugs and they can be very effective. talk about things that are effective for a chronic headache, for chronic pain conditions, that’s much tougher, that’s much harder. so given the amount of unmet need that there’s around for these kind of conditions, I think that’s really one of the main motivating factors for why I want to try to get understanding of why it is that people develop these problems.

Zameel Cader: then not just that, but to then see how do I then make that work into something that’s gonna ultimately benefit people who are suffering with, pain, headache and indeed other neurological problems?

Grant: What makes pain and migraine scientifically challenging?

Zameel Cader: I think almost every single person working on their particular thing will say it’s really challenging, and they’re absolutely right. I think the human body is such a fascinating thing so complex, so many different factors systems working together, trying to keep balance, trying to keep us alive and functional for a good number of years. And it’s extraordinary that it does but of course, it’s not then surprising when things don’t work and things do start to break down or in some way or the other. headache and pain is a good example of that, I think, because we need pain. Pain is absolutely essential for our survival. It tells us about the world around us, what’s harmful, what to avoid. It teaches us things. So it’s a really fundamental aspect of being a person. Being a human is experiencing that.

Zameel Cader: And when people develop a pain disorder, then what is perhaps normal, becomes not normal, and there’s lots of different reasons why that might occur. so you ask the question, why is it so challenging?

Zameel Cader: I think firstly, it’s the causes of why someone develops a pain disorder are really varied different factors, coming together. And again, as with many disorders, it’s a combination of the genes that you inherit that may make you vulnerable to certain things, with the environment that you’re exposed to, and the environment in the very broadest sense.

Zameel Cader: So you’ve got these factors playing together interacting with one another. And then secondly, if you think about how we appreciate pain and we sense pain, then you can recognize that the system really quite complex. So there is a pathway from, for example, if we take think about pain experienced by an injury to our skin.

Zameel Cader: Let’s say you get a burn. There’s a pathway that goes from our skin, travels to the spinal cord, and then from the spinal cord, it travels up to our brain. But along that pathway, there are lots of places where there are interactions that take place, modulation of the signal, and that starts probably right at the skin. of different cell types that are present that are determining how a sensation is interpreted. as it– as the signal gets sent into the spinal cord, there are circuits that are present in the spinal cord modify pain traffic. And alongside the nerve circuit, there are other cell types that are important, which again, can modify how the neurons behave. then we get up to the brain.

Zameel Cader: And then in the brain, hugely complex organ, of course, and lots of different regions in the brain that can then work together be able to, provide context to the pain, for example, an emotional aspect to the pain. And it goes beyond that purely physical phenomena of some kind of stimulus turning into electrical signal, which is then detected into something that’s much more layered with meaning and emotion and so on.

Zameel Cader: And so that’s part of the reason why trying to first understand pain and then to develop treatments against pain

Zameel Cader: can be quite tough.

Grant: When you see patients and then return to the lab, what kinds of questions really stick with you?

Zameel Cader: This is really important. I think for a long time in research knew best or doctors knew best and would pick on things that were, perhaps of academic interest, and a lot of the time that may have coincided with things that were also important for making a difference to patients. But what I’ve been really heartened to see over sort of years is the importance of the patient voice and the carer’s voice in saying, what are the things that are important for us to study? again, as a doctor, what I want to do is to make a difference to the people that I see, to reduce the burden of the pain that they have and other symptoms associated with that, that may vary from one person to the other. And for some people, that might be a particular intensity or type of pain that they experience. For someone else, it may be the unpredictability of their pain.

Zameel Cader: Or for another person, it may be because they get severe nausea or light sensitivity with, let’s say, their headache attack. And understanding what the priorities are for the patients now I think has become increasingly something that I really take into account when I’m then trying to what area of research I should be really trying to work hard on. of course, the thing that drew me into this field in the first place is because I really like trying to understand how things work. What are the fundamental building blocks? So yes, my priorities will be set by patients, and they can often bring incredible insights into how we should run the research. I’ll als- always try to link that back into how do I break it down? How do I think about the that might be working together to then induce that pain state?

Zameel Cader: And then from there, we can then think about how we might develop drugs or other treatments that might help

Zameel Cader: solve that problem.

Grant: So turning that around what do researchers often understand about headache or pain disorders that patients rarely get told clearly?

Zameel Cader: There’s often a gap between, what’s being carried out in the lab and what’s, there in terms of what patients and the people that I see are understanding. But I think often, it’s not so much the concept because I think people with lived experience of pain or people who look after those with pain really very much grasp the kind of scientific concepts are. And in fact, they’re thirsty to be able to understand it. So I think it’s more maybe a difference in perspective and urgency, and I think as a researcher, you realize that things take quite a lot of time. And the process of scientific discovery is often uncertain, and that can be quite difficult to convey, the uncertainty, and that even when a study comes out and it looks as if that is an answer, it’s not really an answer.

Zameel Cader: It’s a current position, which in the future, may be reinforced with further studies or may be dispelled, and that’s an important part of the process. So that uncertainty, I think, is probably something that researchers appreciate and may be a bit more difficult for people who are suffering with these con– conditions to grasp in the same way.

Zameel Cader: So I think it’s probably one important area.

Grant: As a researcher, how do you deconvolve the subjective experience of pain from measurable biology without losing what matters clinically?

Zameel Cader: That is a really tough question, Grant, and I’m not sure that anyone has really worked it out yet. Because one thing is really true in the pain field, which is that we really suck at making drugs work for pain come through our research studies. when we do work in the lab it works beautifully very often. And yet when we get through to the clinic, in the clinical trial or actually being given to patients, there’s a really big between what we see in the lab and what’s there clinically. And I think at least part of it is because of that subjective experience that is such an essential component of pain.

Zameel Cader: And, some of the other things that we’ve just talked, talked about previously, which is pain isn’t just a one-dimensional thing. There’s lots of other factors, and for different people, different things matter. So unless your drug is hitting the things that matter, then a person may not necessarily respond in a clinical trial in the way that you might expect them to respond. So I think there are some things that are possible to be able to take into account of, and there are some things that are really challenging to be able to take into account of. I’m a scientist that tends to work with cell models And trying to encapsulate subjective experience in a cell model is not gonna happen. So often the aspects that I’m working with is simpler

Zameel Cader: aspects

Zameel Cader: in some ways of the biology where we can more strongly the things that we are seeing in the lab with more definitive markers that might be present in patients. And this is a goal, again across not just the pain field, but much of drug discovery, is to try to pick out features that patients with a condition have that can a objective robustness to them, because then I think we’re much more likely to be successful in our efforts to be able to translate from lab to clinic. And we call those kind of things endophenotypes, for example, where there may be a particular feature that a patient exhibits that shows more consistency, or it may be a biomarker that again serves that kind of purpose.

Zameel Cader: And so although I can’t capture subjectivity in the models that I use what I do certainly try to do is to try to build correlations the cell models and what the patient experiences, and that’s, I think, a key goal of the work in my lab. And I apply that across a lot of the work that I do, is to try and establish those types of

Zameel Cader: correlations.

Grant: What makes a headache disorder a good window into broader questions in neuroscience?

Zameel Cader: It’s a fantastic disorder to be thinking about the brain in general because it is something that clearly is affecting multiple brain systems and, the brain hugely fascinating, intricately complicated system. And when someone develops a headache, and I’m gonna particularly stick to migraine, when I’m talking about headache.

Zameel Cader: And the reason I say I’m sticking to migraine is because a lot of the changes that are occurring in the brain a migraine episode have been characterized. We may not understand of the changes that occurs in a migraine, but through various approaches and techniques, we’ve at least been able to get a window into what’s changing. And migraine is a very common headache disorder and perhaps it’s the commonest type of headache disorder that clinicians see. So the commonest headache that people get is tension type headache, but most people put up with that, take a simple painkiller, and it goes away. But if a headache keeps coming back and is debilitating, it’s almost certainly a migraine. And people with migraine a third of them get something called an aura, are transient neurological problems that affect them maybe for about fifteen to thirty minutes, and then it completely resolves.

Zameel Cader: And the commonest type of aura that people experience is a visual aura where they may experience flashing lights, zigzag lines, colors and so on. And so one of the earliest discoveries in the migraine field around what’s occurring in the brain when you’re experiencing an aura. And there is this phenomena called cortical spreading depression, which is a wave of activity that slowly spreads across a brain region.

Grant: What do you think neuroscience drug discovery has historically gotten wrong?

Zameel Cader: So again, as we’ve talked about already the brain is complex. Neurological disorders are complicated and relying upon models that are in some ways simplistic and give easy answers may not be the right approach. Because whilst they can allow us to get through the very earlier stages of drug discovery all of our milestones and success criteria, what’s very clear is that when you meet patients for the first time with your candidate drug, very often it fails. And I think more time spent in what’s called the pre-clinical phase could well increase the chances of success.

Zameel Cader: And two examples of how more careful consideration in the early phases can improve success. The first is the understanding that the incorporation of human genetics significantly increases the likelihood that your drug is gonna be successful, and that’s been shown now, I think, very clearly by the recent successes that if you’ve got a genetic association with your drug target, then that’s much more likely to succeed.

Zameel Cader: So in other words, making the effort to do human genetic studies is worthwhile. The second is that we should be using human cells in order to be able to test compounds, to be able to test whether the targets are relevant. For the longest time, the neuroscience drug discovery community have utilized animal models, and they are very useful because they provide a whole system in order to be able to test your mechanistic hypotheses, for you to be able to test your drug candidates.

Zameel Cader: But the big problem, of course is that they’re not humans. And with the advent of modern molecular studies, single-cell transcriptomics, for example It’s becoming more and more clear just how different at the molecular level the human system is from the mouse, which is one of the commonest animal systems that’s used. That’s true at the molecular level, it’s true at the functional level, and of course, if you take a mouse brain and a human brain, you can see huge, literally, differences between the two, both in terms of the size, but also in terms of the complexity of the of the structure of the brains. And so in many ways, it’s surprising that things that were developed using animal models were successful at all. But I think as, with the low-hanging fruit now taken, we need to fully embrace that we really need much better validation with human tissue.

Zameel Cader: And that’s the field that I work in, which is in human stem cell models.

Zameel Cader: And the reason I work in that area is because getting access to human tissue isn’t always easy or straightforward, particularly nerve tissue or brain tissue. human stem cells provides us a way to be able to get access to that and to be able to experiment on those types of systems.

Grant: So what does a more human-centered model of drug discovery look like in practice? What does that stack look like?

Zameel Cader: I think it’s fair to say that you won’t find a drug discovery pipeline now that doesn’t have human as part of its workflow because limitations of the older um, paradigms are just too apparent. And different companies whether they be pharma or whether they be biotech working in the drug discovery space use different ways of getting access to human-centric approaches. Some companies will have embedded within them, groups for example, develop human stem cell models, and that’s part of their own internal pipeline. But many companies also use external collaborations, and there’s good justification for going to external. One is that working with human cells is non-trivial. It’s technically very demanding Requires quite a lot of time and effort, as well as cost.

Zameel Cader: And so it would be beneficial for many companies to go to someone with an established credibility with working with human cells in a particular area. And there are several options, I think, for these kind of human systems to be incorporated into drug discovery workflows. So you can work with academic groups. And from an academic perspective, I’ve engaged with lots of biotechs and pharma and those have been some of the best research programs that I’ve run. For example, I ran a consortium called StemBANCC, established

Zameel Cader: for

Zameel Cader: drug discovery

Zameel Cader: of the earlier stages of induced pluripotent stem cell, development. And that was fantastic because we worked with companies like Roche, with Johnson & Johnson, Pfizer

Zameel Cader: and

Zameel Cader: so on, to be able to develop these human cellular models for that specific purpose of drug discovery and to make it available for both academic and industry researchers. And such public-private partnerships are ongoing because very often, the material needed to make stem cells is present in hospitals and in academic groups, so one has access to, to that kind of resource, as well as leading-edge expertise in a particular disease. Then on the other end of the spectrum academic groups, there are also contract research organizations can provide cells, human cells or can provide assay services. And, one of the things that I felt a number of years ago was that there was a lack of that kind of expertise and provision for the community.

Zameel Cader: that’s what led to founding a contract research organization called Oxford StemTech which has now been running for four or five years and provides, I think, a much needed expertise in the field for neuroscience human cell assays.

Grant: How do you decide when a human stem cell model is telling you something biologically meaningful rather than just technically impressive?

Zameel Cader: Again, another of challenging question one that the field is constantly wrestling with. There are lots of different types of signals that you can get from human cell models, and the challenge for the researcher is trying to decide whether the signal that they’re getting, is it an artifact? Is it consequence of the platform that one is using? Is there real biology being demonstrated? then is there real biology that’s being demonstrated that’s relevant to the patient condition? So there’s lots of different layers, different levels of findings, phenotypes that you might observe in human cell models. And one of the jobs of the researcher is to try to work out what’s what. And there are different ways of doing that. And of course perhaps one of the important things is to reduce the likelihood that what you’re observing is just technical noise.

Zameel Cader: And you do that by ensuring that you’ve got sufficient replication in your studies, your study is well-designed, and that the statistics that you undertake are appropriate. And again, the research and the analysis is designed in a way that you’re gonna be able to get meaningful and robust answers. And that’s non-trivial, and I think that’s a field that’s continually evolving. And that’s one aspect. The second aspect is, again, thinking about relevance, things that are biologically meaningful and patient-relevant. Again, perhaps coming back to an earlier question that you had, you know, what is it that I think drug discovery companies and those in that field are doing wrong? One of the things that I think we are not embracing enough is diversity of patients and of people when we’re doing our research studies in the lab.

Zameel Cader: And again, this comes down to resources and the challenges of being able to do studies at scale. very often when we do a stem cell study, we might take a stem cell line from perhaps three, four with a particular condition and three, four donors who don’t have a condition. But that’s just a very small group, subgroup of the population that you might be interested in studying.

Zameel Cader: So you’re really not able to model the diversity that’s present for that condition. So I’m really a firm believer of trying to increase the numbers of donor lines that we interrogate when we do our research studies. And I liken it to how we used to do genetic studies or genomic studies 15, 20 years ago. resources and cost limitations meant that we often did candidate gene studies rather than approaching in an agnostic way, and we would, you know, take maybe a gene, take a few handful of individuals with a condition or without a condition, and test for association of that gene. And that might produce an association, but almost inevitably, when the large scale studies came along, they, they turned out to be false associations. I worry that we’re in a similar phase with cell studies, that we do small scale cell studies at the moment.

Zameel Cader: We find these associations, but when we get around to doing the larger scale stuff, which we will do, much of those are gonna really stand the test of time? I’m worried that many won’t.

Grant: Major concern. So what is the right way to think about the gap between a cell model, a patient, and the treatment?

Zameel Cader: I don’t think there is one right way. I think there are multiple different approaches. So, Again, if I stick to my field of human stem cell models, one of the most amazing things about working with human induced pluripotent stem cells is that they’re capturing the genetics of the individual. Because what you do is you take a blood cell, for example, from an adult with a condition, and from the blood cells, and specifically a common starting cell is an is a precursor to a red blood cell, which still has its nucleus and still has its DNA, and that’s turned into a stem cell. So the DNA that was present in that person is now in your stem cell. So all of the genetics and the risk factors that have predisposed that individual to a condition that they may develop or have are present in your model. So that already brings a gap closer between your model the person.

Zameel Cader: Now what’s missing is the environmental exposure And that is something that we can reintroduce. And I think that’s often something that people appreciate with the cell models. V-very often it’s all about the genetics, it’s all about the fact that they capture the genetic susceptibility. with our cell models, we also have an opportunity to be able to modify environment that that cell experiences over quite prolonged periods of time. So for example, we did that when we studied neurodevelopmental processes for people with a type of epilepsy And what we found was that cells that were carrying a mutation in a fundamental gene that controlled cell growth controlled the balance, the energy balance. so it’s a gene that’s really important for almost every cell, a mutation you’d imagine would be absolutely disabling and deleterious.

Zameel Cader: And what we found was that by modifying the amount of glucose that was present and being delivered to those cells, either completely rescued any abnormalities that were there magnified the abnormalities. And so, that just for me really reinforces the importance of gene environment interactions. And when we look at our cell models, just be thinking of the genes.

Zameel Cader: We need to be thinking about what the environment that these cells should be being exposed to that might then allow the disease to manifest. Now, coming back to pain, the question is: are the challenges that a cell might need to be exposed to to be able to induce a pain-like state? And we don’t know. But we’ve got some good candidates. So that might, for example, include adding inflammatory factors, cytokines, and interleukins to our pain nerves might then lead the pain nerves to enter a kind of sensitized state. So this is one of the things that we’re gonna be working on over the coming years, is to try to understand that gene-environment interaction to then be able to bring the patient the cell models closer together. Then your other question was around treatment, and then how do we get to treatment?

Zameel Cader: So our approach is, again, to go back to the patient and to select patients that we’re going to model based upon how they responded to treatments in life. So you might imagine someone with migraine, and what we have are individuals who’ve responded to one of the new migraine drugs. Called the gepants. And so patients who have responded, we can make their cell models, and the patients who haven’t responded, we can make their cell models. And it’s a question that we haven’t yet answered, but we’re hoping to answer is, do their cell models from the people who respond to a treatment, are they different to the ones who don’t respond to a treatment?

Zameel Cader: So that’s how you might then be able to bridge across two treatments.

Grant: What are the hardest parts of building human models of the blood-brain barrier?

Zameel Cader: So the blood-brain barrier is a really interesting structure. It’s made of brain endothelial cells, and these are specialized cells specific to the brain because they have what are called tight junctions between them. That means most substances can’t cross from the blood vessel side into the brain. They also, these endothelial cells, have very dampened transport mechanisms. And what that means is that when a substance might enter an endothelial cell itself, in other tissues, normally that substance might then get ferried across the endothelial cells and then cross the membrane on the other side to get into the tissue. But in the brain, that is dampened down. That’s a process called transcytosis. That’s dampened down. So we have that. That’s, that’s the core of the blood-brain barrier.

Zameel Cader: But in addition to the blood-brain barrier, we also have other cell types that are really essential for barrier function, and one of those other cell types is called pericytes. And we know that they’re important because if you remove pericytes experimentally, the blood-brain barrier becomes leaky. So we know that they have a very important role. Then another really important cell type are astrocytes, and astrocytes are on the brain side, they contact the vascular cells, the endothelial cells, and the pericytes. And again, they’re also really important for regulating barrier function. And we also think that other immune cells like microglia probably have an important role as well. So there is the anatomy of the blood-brain barrier. There’s a polarization. You’ve got certain features on one side and other features on the other.

Zameel Cader: Then you’ve got a cellular complexity, and you also have a functional coupling between the different cells. And so trying to reproduce that in vitro quite challenging. And there are models that have been reported, but I think there is still more work to be done to really to be able to establish a bonafide blood-brain barrier model.

Grant: How has human genetics changed the way you think about neurological disease?

Zameel Cader: I started off in my sort of research journey studying human genetics. After I qualified for medical school and did the early training, I then started a PhD. And my PhD was trying to identify a gene causing a Mendelian disorder where patients would suffer episodes of imbalance and that condition’s called episodic ataxia. by Mendelian disorder, what I mean by that is, is that a mutation in a single gene is sufficient to cause the condition. And we were gene hunting essentially at that time for rare conditions that were being caused by and abnormalities in single genes. And this was 25 years ago, perhaps. and since then, the field of human genetics and genomics has really changed, transformed into something that certainly would’ve been unrecognizable to me when I was doing genetics back in my PhD.

Zameel Cader: And we’ve moved from the study of these rare single gene mutations to now complex disease, where it’s a combination of both genes and the environment that cause disease, and where the genes are no longer single genes, but small changes, variants in the genome very slightly increase your risk on their own and collectively perhaps increase your susceptibility to a condition. And this whole field of human genetics and genomics has been transformative, I think, across all of medical science because taught us it’s shown us many of the processes and mechanisms that might underlie the common diseases that humans suffer. It’s also illuminated mechanisms that might be targeted for drugs. So it’s part of the research that I do. It’s a fundamental part of the research that I do. You can’t ignore genetics because it’s such a strong factor in why we develop disease.

Zameel Cader: And as we’ve talked about already, this, it comes back to the basis of our cell models and how we drive drug discovery. I think that there is still a lot to be done in this space. Although there are now many genetic studies genome-wide association studies where we’ve been able to identify variants associated with a condition, it is still very much the case that for many of these variants, we really don’t know what they do. is it that you go from a variant? How does that variant increase your risk of disease? How do variants in one gene or one part of your genome interact with others? I would hope that some of the cell models that we generate some of the research that we do can help, better understand kind of mechanisms.

Grant: How did your path from medicine into genetics and neuroscience unfold?

Zameel Cader: I was an unusual person perhaps in that I always knew, even when I started medical school, that I wanted to pursue medical research. And I had a first proper exposure to research at the University of Birmingham and they do an intercalated medical degree. And that’s done in the third year of our medicine program. And so I did a degree in pharmacology, and I was in a department which was working on long-term potentiation, which is a form of cellular memory. And it was also around the time that a molecule called G proteins the diversity of G proteins were just being discovered. And it was just absolutely fascinating. I said I knew I wanted to do research, but that absolutely affirmed that, that, that was the career that I absolutely wanted, and reinforced just how, incredible neuroscience is. And that cemented my ambitions.

Zameel Cader: So I then finished my medical degree, came to Oxford shortly afterwards very quickly entered my PhD going into genetics and genomics. Once I’d finished my PhD and then my medical training I then was successful in getting a, what’s called a Medical Research Council Clinician Scientist Fellowship, which allowed me to establish my own independent research career. And then from then on, it’s really been trying to continue research in the way that I’ve described.

Grant: What have you learned from leading large collaborations that you could not have learned running a single lab?

Zameel Cader: They’re very different beasts running large collaborations and running labs. running large collaborations requires, really, it needs a passion working with others to solve problems, communicating and try to find the best path with the talent that you’ve got available.

Zameel Cader: Because, it’s joyous and it’s incredible to be working with people at the forefront of their particular research area, being able to pull people together on a problem to solve. And you get access

Zameel Cader: to

Zameel Cader: expertise and resources,

Zameel Cader: which

Zameel Cader: isn’t so easy when you’re running your own individual lab. Now, of course, when you’re running your own individual lab, you’re often looking to establish collaborations, and you work hard to do that. But the mindset and the approach and the sort of immediacy of running multi-group research experiments is very different when you’re running a consortium.

Grant: What surprised you the most about building a company?

Zameel Cader: So building a company has its rewards and its challenges, and it’s really tough, and gives back a lot. It’s just so many different emotions that you go through when you’re running a company. I was very fortunate to have two exceptional co-founders for Oxford StemTech who were previously researchers in my lab then left to start the company up. So we started off on a shoestring and have slowly been building the company over the last several years. and it’s a fantastic team. And I was not sure… I started off in academia. I was never sure whether I would get the same rewards, stimulation and output that I had in academia from a company. And I’ve really been pleasantly surprised just how fantastic the team at Oxford StemTech are.

Zameel Cader: They have been able to establish workflows, experiments, run internal R&D at pace, and to be able to deliver robust, reliable results, generating actually some amazing new insights into mechanisms along the way.

Zameel Cader: Even though our priorities are quite different, an academic lab versus a company, nevertheless, some of the things that we’ve been finding in a company have just been extraordinary. and one of the things that I think perhaps benefits the company from having an academic co-founder is that I’m still keen to publish, whether it’s a company or whether it’s in academia. So I hope some of the things that we’ve been finding in sort of recent years in Oxford StemTech will be published soon. And I would certainly say that you have intellectual drive fascinating outputs, whether you’re in an academic lab or whether you’re in a company

Grant: What advice would you give to early career scientists who want to work on translational problems without losing scientific depth?

Zameel Cader: The main factor for a successful career in translational science is motivation, passion, and those come from understanding that there is unmet need. People are suffering with conditions that if we get our science right, we can make a real difference to. And if you start from that position, you’re driven by your patients, you’re driven by the need that’s out there, then you will want to ensure that the science that you do is robust and is meaningful. And that necessarily means that it has to have scientific depth because doing things half-baked to get a quick win, to get a quick result without showing reproducibility isn’t going to help your patients. It’s just gonna generate more uncertainty, might undermine other work that’s going on.

Zameel Cader: it’s really important that you maintain integrity, that you train yourself in, scientific thinking, and certainly with people that come through my lab or through the company, think that ability to think critically, to think deeply, really focus on the question that you’re trying to answer and not to be, led by sparkly new science but really to get back to the core of the problem.

Zameel Cader: That’s really what you want to drive you. That’s what I try to impart now, I think, to people that come through that I, have the privilege of being mentor to or supervising or line managing and to develop those skills because, whether you learn this technique or that technique or whether you get this paper or that paper, that’s not what’s gonna make you successful.

Zameel Cader: What’s gonna make you successful is that… those core skills that we’ve discussed.

Grant: And what do you realistically hope the field will be able to do for patients 10 years from now that it can’t do today?

Zameel Cader: So coming back to the translational sort of science, I think we are lacking meaningful treatments for chronic pain conditions. The treatment landscape for headache is much, much better. It’s been transformed by the recent therapies that are available, and I think chronic pain could follow. So I would hope in 10 years we would have treatments like we have for migraine and headache, and I think that’s achievable.

Grant: Zam, thank you so much. It’s been a great conversation.

Zameel Cader: Great. Thanks so much, Grant.

The Bioinformatics CRO Podcast

Episode 90 with Adam Woolfe

Dr. Adam Woolfe, computational biologist and bioinformatics R&D leader at Bio-Rad, discusses building bioinformatics tools for scientific instruments and new approaches to antibody discovery.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Adam Woolfe

Adam Woolfe is a computational biologist and bioinformatics R&D leader at Bio-Rad, where he leads bioinformatics for the Saber project.

Transcript of Episode 90: Adam Woolfe

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. I’m your host, Grant Belgard. Today I’m joined by Dr. Adam Woolfe, a computational biologist and bioinformatics R&D leader at Bio-Rad. Adam’s career has taken him from comparative genomics and gene regulation to single cell technologies, antibody discovery, and computational immunology. His recent work includes Paraplume, a sequence-based approach to predicting antibody binding regions using protein language models. We’ll discuss what he’s working on today, the path that brought him there, and what he has learned about building up career at the intersection of biology and computation. Adam, welcome to the podcast.

Adam Woolfe: Nice to see you again, Grant. Thanks for inviting me here.

Grant Belgard: So for listeners meeting you for the first time, how would you describe your current role and the questions you’re responsible for?

Adam Woolfe: Okay, so my role is basically to develop the bioinformatics tools pipelines and software necessary to develop our instruments and also to support the customers that use our instruments. So I think your listeners will be familiar with Bio-Rad. I think it’s it’s one of the biggest sellers of scientific instruments. Things like digital PCR, PCR thermocyclers, gel electrophoresis machines. If you have worked in a lab, and I guess a lot of the people who do bioinformatics have not. But if you… the people who make your data will probably know about Bio-Rad because it is such a big company. So, I’m based in Paris, France, or just outside Paris, France, and how I ended up here was basically that I used to work in a startup company called Saber Bio. So I’m now part of the Saber Bio project in Bio-Rad.

Adam Woolfe: But Saber Bio was a startup company a couple of years ago, and we actually got acquired by Bio-Rad a couple of years ago to bring in single, like a droplet microfluidic single-cell science into the Bio-Rad portfolio. So that’s what we were doing. We were developing a technology called droplet microfluidics, which is basically the thing that drives a lot of single-cell science today. People may be familiar with uh, 10x te- uh, Chromium machine, to do the single-cell science, although there are a number of other technologies. But theirs is based on droplet microfluidics, and basically you’re creating picoliter, like super small droplets that you flow through channels in a chip. And what what we do differently compared to other, o- other people is that we have expertise in in-droplet assays.

Adam Woolfe: So we actually assay the function of the cell in the droplet using techniques to understand what’s going in the droplet, and then we can use droplet microfluidics to sort the droplet down channels depending on the function that they have. So we apply this, of course with the, the machine that we’re developing actually is applied to identifying therapeutic immunoglobulins, so antibodies and TCRs and what we do is we can assay the cells that produce those things, the B cells and the T cells for the function that we’re interested in. So for instance, for the antibodies, we’re interested in binding, whether they bind something or maybe whether they, whether they have a particular function they can internalize or they can activate a cell. And all that can be done inside that little droplet using fluorescent signals and laser lines. And we can sort those cells.

Adam Woolfe: We can enrich for a population of cells that are producing molecules that we’re interested in. And this is applied obviously to therapeutics, right? So antibodies make up, one of the most successful th- range of therapeutics, biologics, right? That that are on the market. They… the value of antibody biologics is something like three hundred and fifty billion dollars today, that could, could raise to one trillion in in five to 10 years. In fact, out of the five of the top 10 biggest selling drugs, five of them are antibody therapeutics. In fact the top drug which is a, a, an antibody called Keytruda made by Merck basically brings in thirty-two billion dollars to Merck every single year. So these things are big business, right? So identifying antibody therapeutics is a big deal these days. And so that’s what our instruments… they don’t necessarily have to be applied to that.

Adam Woolfe: You can use single-cell science to understand other things as well. We work with groups looking at autoimmunity or whatever. But basically, if you… One of the biggest applications obviously value-wise, is to identify new antibody drugs.

Grant Belgard: What parts of your contributions at Bio-Rad are least obvious from your job title?

Adam Woolfe: So certainly in a big company, you have to sell yourself to… Or you sell your ideas and the, the things that you wanna do, you have to sell it to the higher management or to non-scientists even to marketing team. In fact, the marketing team is like one of the most important aspects in a big company. Like what is gonna bring you value? You can, as a scientist, you often think, oh, this, the, the most interesting things or the most novel things are gonna be the things that will propel you forward. But actually, in a big company where, the bottom line is whether you can make money from something, right? Is more important. So a part of my job is to explain the value of a bioinformatics function or a bioinformatics tool that will facilitate both the analysis of data coming from the instruments that we’re creating and bring value and money to the company.

Grant Belgard: Can you trace a representative project from biological question through data generation and analysis to a final decision?

Adam Woolfe: As I mentioned I spend some of my time on R&D to to help people prioritize antibodies. Okay? So when someone uses an instrument, they may end up with, several hundred candidates that bind their antigen of interest or whatever. But the, the next stage is which of those should I take forward, to to the next stage of testing and validation, which is expensive, right? Developing drugs is expensive. If we can spend time before if the person can prioritize computationally the candidates that they can take forward, they can save a lot of time and money, later on further down the road. Because antibodies, as I say, are expensive to develop, and they fail very often because they have to be produced as drugs, and that process is quite rigorous.

Adam Woolfe: So they have to be produced at high concentration high pH high temperature, and that process can cause antibodies to aggregate and to lose their function or whatever. And so even if in the initial stages your antibody of interest was super good binder it doesn’t mean that it’s gonna make a good drug, right? It can fail in many different ways. And so computational prioritization is a really huge field that is actively worked on by many people in the world today because of how much money you can save downstream. And so what we wanted to do is to get a handle on the fundamental function of the antibody as it were. Okay, so what is the fundamental function of the antibody?

Adam Woolfe: It is to bind something in a very specific way, what’s called the affinity of the antibody, and the, the active components of that are literally the amino acids on the outside of the antibody that interact directly with the antigens, right? So those are the specific interactions between the antibody and its cognate antigen. Those are the active components. Now, there are other parts obviously, of the antibody that are important for ensuring that those amino acids are in the correct position, that they fold correctly and that the whole molecule is correct. But active parts are really the thing that drive the affinity. So if we can identify those, that would be good. Now, there were a number of different approaches that were already out there in the field but they often were quite their accuracy was not great, and they often relied on structure, right? So structure is fantastic.

Adam Woolfe: It, it’s revolutionized the field, the ability to in silico fold up a protein and know what it looks like in three dimensions, of course, is an amazing thing, right? AlphaFold has just revolutionized the field, and there are many other, and there are specific approaches to folding antibodies like what AlphaFold brought. But that obviously that comes with a lot of computational resources that require you, to expend to just to get the result. And that… And if you wanna do that on a repertoire level, so a repertoire, an antibody repertoire is a group of antibodies which can be in the scale of hundreds or thousands or tens of thousands of antibodies that come from a specific immu- immunization or immune event from an organism that you need to characterize, right? So if you have a very computationally expensive, process, you can only do stuff on a one-by-one basis.

Adam Woolfe: And this is true, in anything in bioinformatics. If something takes 20 minutes per molecule to run, on a GPU okay, that’s not gonna be very useful. It’s gonna be useful on… to characterize maybe one or two sequences. But if you need to look at this whole process in a kind of repertoire level, and, you can gain insights when you do high throughput, that you can’t gain when you’re looking at things on a one, on a single level. we wanted to have the ability to look at these specific residues that are so important and look at them in a very high throughput way. So we wanted to make sure that The process that we the, the model and the tool that we created was quick and was accurate, and it was at least as accurate as what was already available in the field.

Adam Woolfe: So we said, “Okay, let’s, there are these new things called protein language models, which are foundational models, which basically encapsulate a lot of evolutionary and structural and functional information in them because they’re trained on millions or billions in the case of antibody sequences. Sequence so they, already understand the fundamentals of what it means to be an antibody, what it means to be a protein before you fine-tune them on things and say, “Okay, now you understand about proteins, now you understand about antibodies. want you to learn what it means to be a paratope,” right? So now we trained those models on a small relatively small number of sequences because The, the availability of good, high-quality data in the field of immunoinformatics is actually quite limited. So you’ll come across this problem all the time.

Adam Woolfe: There are just not that much data available that you can train your model on. So the using these protein language models is pretty useful in that you don’t have to give it the information from scratch. It already knows the fundamentals already, and then you just have to tell it, “Okay, these are the things that look like this.” So we used structure, 3D structures of real antibody-antigen interactions to train the model and see whether we could get a good accuracy. And of course the novelty was that we didn’t just use one protein language model because there are a whole bunch of different protein language models that are trained on different datasets. Some are on just on proteins, some are just on antibodies, some are subset of antibodies, whatever. They all have different types of information encoded in them and no one model is necessarily the one model that should be used to train this.

Adam Woolfe: So we actually, we… What we did was we just stacked a whole bunch of them together the embeddings of all these across the across the sequences, and then saw whether that would improve the accuracy. And in fact it did. And using embeddings is pretty quick. So we were now able not to run one sequence in a minute, but like hundreds of sequences or a thousand sequences in a few minutes. So now we can cover the repertoire in a very rapid manner in a way that we just couldn’t before. And now trends that are interesting in that field like, like just understanding what… how do paratopes change when an antibody becomes exposed to an antigen, right? In the body when it goes from a naive to an antigen-specific antibody, how does the paratope, which is the active, component of that, how does it change? And now, Once we created that model, we were able to answer those kind of questions.

Grant Belgard: What are the most consequential false positives and false negatives in antibody or immune cell discovery?

Adam Woolfe: Okay, so if you are using computational means to predict particular functions, obviously the consequences can be great. As I mentioned, the, the advantages obviously are that you can prioritize things quickly, and you can choose antibodies that will likely not fail when you turn them into drugs. Now, the consequences of a false positive are that, you can, predict something that looks good. It may look like on paper that it works very well, but actually when you put it in the lab and you express it and you whatever it causes problems because, your model just wasn’t enough. Now, the false negative obviously is that you can miss things that were good, right? Those rare events that look possibly like a false positive when in fact it’s not a false positive at all.

Adam Woolfe: And you can, yeah, you can miss out on those, those things that would have made a fantastic drug just because the models or the predictions told you that it was not gonna be great. So yeah, we still have a long way to go. Predictions are not, biblical, right? They’re just informed guesses, right? That you just have to take with a bit of pinch of salt and some people maybe take them maybe too far. But I think there is a certain level of skepticism, I think at least in the experimental community of AI models, which I think is completely justified. But I think it’s– if you only have the ability to test a small number of sequences ’cause your budget, it’s very expensive to, to, to do that, then I think prioritization is still a pretty good way to go.

Grant Belgard: How do you approach computational workflow design when an assay or instrument’s still changing?

Adam Woolfe: Ah. So this is something that I’ve spent literally the last 12 years doing. I was in originally in a company called HiFi Bio, where we were also developing droplet microfluidics screening for antibody discovery. We were doing it for internal screening rather than trying to create a machine that was used by people. But you have to go through so many cycles of changing the approaches from scratch. It just… yeah. So effectively, you have to just be flexible and modular in your approaches so that you can take out, certain parts and just modify them in ways that will allow you to deal with the new approaches. I think just flexibility in your software design, in your analysis approaches that can, that you can switch in and out of stuff that, so that you don’t have to reinvent the wheel, but You certainly don’t have to, spend a lot of time changing stuff.

Grant Belgard: What does product readiness demand that publication readiness does not?

Adam Woolfe: So publication you have to be right once or a, a few times. Your experiment has to be, right and robust enough to get through review. But it doesn’t necessarily have to be, true in every context, right? So I think there was a study done a few years ago that showed that something like 75% of published experiments were not reproducible, right? So it just shows that the requirement for reproducibility is something that is not generally done generally because it takes time and money and effort. And once people have the result, they just wanna get out there and publish it as quickly as possible, right? So they don’t think about the reproducibility.

Adam Woolfe: So when you’re working in a company reproducibility, robustness, and the fact that a machine has to be taken and used by different people in different contexts in different environments, and it has to work every single time, is just a completely different ballgame that you have to deal with. So yeah, the two worlds, I think couldn’t be much more different. And in fact, know, a lot of companies are based on academic research that was done that looked promising, and they’re like, “Yeah, you know, just click your fingers,” and all we have to do is just develop this into a machine or into an approach, and it’ll work first time. And boy, is that wrong. The ideas are great, but turning those proof of concept experiments that perhaps drove a, biotech start-up were, uh… it takes a lot more to actually make it into a functioning product because it’s just not as easy as people think.

Grant Belgard: What evidence makes you trust or distrust a model enough to let it influence an experiment?

Adam Woolfe: Yeah, this is a really hot topic, I think, in the whole AI community right now, at least in science. I think if your model generalizes outside of the training milieu, like if it works on something that it didn’t really ever come across before, that’s the kind of the golden answer to whether your model is worth doing. Like, Is a lot of cases, I think, where data poisoning occurs. So people are not necessarily aware that the, the data that they’re validating on can look very similar to the data that they’re training on. So if your test data has things that, have high sequence identity or they are of sa-same function they may come through in the test data and of course, the model just says, “Oh, I’ve seen that before, that’s easy.” So it elevates the model’s accuracy rates. It looks like the model is much more accurate than it actually is. And when it…

Adam Woolfe: In fact, when you test the model on data that looks completely different to the model it often fails quite badly. So basically, if a model works well on data it’s never seen before, I think it’s trustworthy. Otherwise, not so much.

Grant Belgard: What’s the hardest handoff among experimental biology, microfluidics, instrumentation, software, and data science? Where do you, uh

Adam Woolfe: yeah. So the, I think the handover, Like I work with I work with engineers, I work with cell biologists, I work with molecular biologists who are working on the instruments, to make them work. And the handover of the data is trying to understand like the, the biases or the problems that propagate silently, from the instrument itself to the data that you’re looking at. It’s like you don’t necessarily have hands-on understanding of the instrument. It’s hard enough to understand bioinformatics, never mind engineering and whatever.

Adam Woolfe: So spending a little bit of time in the lab for sure goes a long way to try to understand the ins and outs of your data, like why you’re seeing what you are seeing and trying to understand that better so that you can mitigate the, the effects of those very specific aspects that you find in your data that are associated with the instrument that was creating it or the cell or molecular biology approach that was used to create the data. I think that’s the most important aspect to understand. Yeah.

Grant Belgard: What changes when a development stage platform becomes part of a larger organization and what should remain unchanged?

Adam Woolfe: Okay, so I’ve been through this process a lot. I’ve been in startup companies with a lot of change. And then as, as I say our startup company was acquired two years ago, so going through that process you understand that the big organizations just work very differently because their aims are very different. Like, when you’re in a startup company it’s effectively proof of concept largely. You have to work fast. You have to be prepared to fail, and you can… And it’s okay because you can… The cycles of development are much faster, right? So you can… You don’t have to talk to someone, you don’t have to sell it. You just have to say, “Okay, look, this is what happened. I think we should try this approach or, improve this,” whatever. And then the, the people go back in the lab and do it and so those cycles go fast.

Adam Woolfe: When you join a big organization so first of all, the teams that you work with expand, right? So you’re dealing with people in a completely different… In a different country. Now you’re dealing with people who are saying, “Okay we need to get this product out on the market in, in, in two years and it needs to be like this and this.” Things have to be more robust. Things have to be more reproducible. Things have to be just… It’s just a very different way of working. Sometimes big organizations, especially if they’ve taken over, or if they’ve acquired a small biotech company it can be like it can destroy the company. Like making that transition, if it… Especially if the team worked really well in that kind of dynamic environment and suddenly when they have to work in a very much more formalized way, it kills something.

Adam Woolfe: It kills something in the motivation, it kills something in the dynamics of the group. And the company realizes it and they go, “Okay, actually, we wanna go back to the way that you were before because we thought you worked really well.” So often sometimes big organizations will say, actually “Don’t worry too much about the formalities that we have,” especially in technologies that they’re not used to, that they haven’t developed these kind of things before. They don’t have templates, right? They don’t have things to say, “Okay, this is the way that we did it before and this…” Because your technology is like something completely new to our organization. So actually let’s… We’re gonna just let you get on with it and because it worked pretty well while you were a startup. Yeah, I think that generally is the conversation that’s had in these, in, in big companies.

Grant Belgard: And now a, a question that uh, is much discussed in our field. How do you expect AI will be impacting our field in the years to come?

Adam Woolfe: There is already the big sort of hot topic in, in, in antibody discovery and development is obviously de novo antibody generation. So now not using screening technologies or phage display or hybridoma or whatever it is that people used to use to find their target is now using computers to actually create antibodies de novo, right? So using generative techniques to create an antibody that’s never been seen before. And then using 3D modeling and other such techniques to try to predict which ones are going to be successful because as you can imagine, the search space in antibodies is absolutely huge, like phenomenal search space. So actually producing, something generatively that, that might actually bind is a huge challenge. And now there are a bunch of different companies today that claim to be able to do this pretty successfully.

Adam Woolfe: there’s companies like Chai-2 Nabla and, all sorts of other ones. Now, they have… They produced white papers that look pretty impressive. But no one has actually seen the data because I guess it’s so sensitive. Like the, the field is so competitive that they don’t want people to see exactly how they did what they did. I think basically as far as from what I hear they basically throw the kitchen sink at AI models with as much interaction data as they can so that the model can basically then predict interactions. But antibodies are slippery. They’re much more challenging than protein-protein interactions just because they’re, the active components of an antibody, which are the basically their complementary determining regions, these are loops that come out of the antibody that interact with the, the antigen are quite dynamic, right?

Adam Woolfe: So they move around a lot, and they can change their shape like after… before and after binding. So they’re much more difficult to predict than general protein-protein interactions. So it’s still a very challenging field, and I think… But AI seems to be every week, every month, there’s new advances in this field and yeah, the golden goose is to be able to produce to predict an antibody very quickly just de novo.

Grant Belgard: Where did your interest in computational biology begin?

Adam Woolfe: I grew up loving computers. I used to do a little bit of copying magazines, you know, uh, like, code. This was, like, in the 1980s, right? I’m pretty old. So I really loved playing around with computers and… But I come from a very scientific family. Both my parents are PhDs, right? So science was quite a strong driving force in my career. And I never really thought about computers as a career. I always thought, there’s science and then there’s computers, and computers are for games or I don’t know, software, whatever, and science is science, and I didn’t really ever think about the two coming together. And it was only really during my… So I did a degree in molecular biology at Manchester University, and as part of that, I spent a year in a lab at Hadassah Medical School in Jerusalem.

Adam Woolfe: I remember one day, one, one of my colleagues coming in and and he was working on some bacterial protein and trying to understand something about why it was working. I can’t remember exactly what it was. But he then said, “Look I’ve threaded my protein onto a, a 3D structure, and it looks like it seems to be this kind of structure.” And I was like, “Wow!” That really, that blew my mind. Wow, you can actually take sequences and computationally predict stuff. That was my first ever exposure to that. And at the end of my degree I took, there were some courses in bioinformatics, and it just happened to be that Manchester was one of the first places in Europe to offer a master’s in bioinformatics, and this was back in 2000, right? So I was o- I was in the right place by the right time ’cause they were offering a fully funded master’s. You even got some money to help you for accommodation.

Adam Woolfe: That’s unheard of now, right?

Grant Belgard: Yeah, it’s nice in the UK.

Adam Woolfe: so it’s- Free so they were offering that and I was like, “Wow, this is amazing. I can put, I can combine computers with biology and, the two passions, you know, together, and that’s just a fant- a fantastic opportunity.” So I did that master’s and then at that it’s funny how the, how these cycles happen because at that time, people were talking about bioinformatics being this this new hot topic, right? And that people were be- were, people were being dragged off the course to get jobs, right? Companies were so desperate for bioinformatics, bioinformaticians, that they were just dragging people off the bioinformatics course before they even finished, giving them high-paying jobs and it’s like, wow, this is gonna be amazing.

Adam Woolfe: And I spent a summer working in a British biotech um, which was an, an old biotech company that was, um- That was quite well known at the time, and I said, “Oh, I’ll finish writing my thesis, and then I’ll look for a job after.” ‘Cause clearly I’ll just find a job very simple. I don’t have to, I don’t have to worry about that. And I graduated two weeks after September 11th the economy collapsed. Bioinformatics companies at that stage the promise of bioinformatics just hadn’t been realized. There was a lot of hype, there, there was a hype cycle, and suddenly people were saying, “Eh.” And it just was a little bit too early, right? The tools were not there. The data was not there. It just it was just too early on and people were, had, had high expectations of this new field called bioinformatics, and it just wasn’t fulfilling. And so a lot of these companies were folding and all their…

Adam Woolfe: the people who had experience in bioinformatics were suddenly on the market. And I was like, “Oh, I can’t compete with my, my, my limited experience in bioinformatics.” So I was unemployed for a year what I kept hearing was, “Oh, people have got PhD,” no. I was like at that stage, I didn’t really think about continuing academia. I’d already spent four years studying, sorry, four years doing molecular biology, one year doing masters of bioinformatics. So five years already university, I was like, “Enough. I, I wanna do something else.” And you have a vision of your life and it doesn’t always work out the way you think. And actually uh, decided to go back and do a PhD because that was pretty much my only choice. And I basically was in the right place at the right time. I’m super privileged that I ended up studying the human genome just after it had been completed.

Adam Woolfe: Like you, you know, I was given the opportunity to study the human genome and apply new bioinformatics techniques, multiple alignment, whole genome multiple alignment using BLAST, whatever, just looking at comparative genomics and finding super interesting things about the human genome where that was completely uncharacterized, right? You stuck a pin and you find something interesting. So I was just in the right place at the right time, and I was, I’m I feel very privileged to have done that. Yeah.

Grant Belgard: And how did single cell biology and immunology enter your career? What made them compelling enough to to stick with it?

Adam Woolfe: Yeah. They actually came in at exactly the same time. I, I hated immunology at university. I found it so boring. There’s this cell, and there’s this cell, and there’s this receptor, and na na, CD4. And you’re just like, “Oh, this is just so boring.” And it was not on my radar at all. And as I was coming toward the end of my second postdoc, I was at the Institut Curie in Paris. This is how I ended up in Paris, okay? This is how, how a lot of scientists end up in countries that they may- maybe never imagined they, they would be in or that’s a very common experience of scientists I think outside of the US. US scientists tend to stay in the US, but scientists outside the US kind of move around. So I was finishing my postdoc, and I was wondering whether to stay in academia, but I wasn’t I wasn’t sure and hadn’t really thought about maybe… I hadn’t thought about industry either.

Adam Woolfe: I was like, I was still pretty wedded to the idea of an academic career ’cause I’d been, I’d had really fantastic experiences in that. But this, this startup company was basically wanted to pair with a, with our lab to look at single-cell chromatin, and we were the experts in chromatin at that point. And so I went to meet with them, and we discussed whatever, and I gave them my opinion and what I thought about the data. And then after the meeting, the guy’s “Oh, know, we’re looking for a bioinformatician to, to to set up the bioinformatics of our company. Would you be interested?” And I was like, “Yeah, that sounds amazing.” So this company, HiFI Bio, was melding micro droplet microfluidics to single-cell science and immunology. They’re trying to find antibody discovery all at the same time. So it’s mel- melding that, those two worlds in a single technology, and it just…

Adam Woolfe: That idea that you could use a technology to learn many different things. You can apply it to many different things. Also to TCR discovery, so looking at what, what you can do with technology that can assay like function at high speed just was very compelling. And in fact what we were– what we started working on in later on when HiFi Bio became a therapeutics company the field of immuno-oncology, right? So this is what really bit me, like the, the ability of the immune system to actually fight cancer, to actually… you can use your own immune system to fight the cancer rather than using an external drug, right? And I remember a couple of years into that my mother-in-law actually got metastatic melanoma. She had– she was diagnosed with metastatic melanoma, and there was buildup of a tumor on her, I think it was her kidneys or liver. And it was pretty fast.

Adam Woolfe: And they put her on this PD-1, immuno-oncology um, treatment. And lucky enough, because actually only thirty to forty percent of people who actually undergo this can actually respond. But lucky enough, she responded and her tumor disappeared. And I was like, “That is amazing. This is amazing.” And that really seeing the effects of something that you, that your technology can, and your company can do is really compelling.

Grant Belgard: What’s a career decision that felt least obvious at the time but became especially consequential?

Adam Woolfe: I think it’s that transition from academia to startup company. As I say, like when you’re in academia you can really, feel like, oh, this is the be-all and end-all. And I think it… My experiences were that what was happening in academia was far more interesting and radical than what was going on in companies. Companies were, for me, it felt like a few years behind the forefronts of science that were being pushed by academic circles. And I think that really was true in the early 2000s, especially with the Human Genome Project and everything else. Really academic science was really pushing the boundaries. at some point I think those advances were now being pushed by, startup companies, right? That like really incredibly interesting science was being done by those companies. And as I say, I wasn’t looking to, to transition into into industry, but it…

Adam Woolfe: And then just, it’s just for me, it just happened, and I was like, “Oh yeah, this seem, this seems pretty cool.” And I made that transition, but it was only afterwards that I realized actually I can make a lot more impact in a world in which the things that I do are actually translated into real drugs, real effects on patients rather than in academia where things are, you make advances in science, but someone ultimately will build on that to do, to help the people that you help. So you… Yeah, in both ways, but I think that transition was not so obvious at the time and now is much more obvious.

Adam Woolfe: And in fact, if I look at all the kind of big names that were in the academic field when I was in academia and now actually in companies, most of them have transitioned out of academia ’cause I think they probably can see the same thing as I did, that a lot of exciting stuff is being done in companies now rather than academia.

Grant Belgard: What should an early career bioinformatician learn deeply today and what can safely be learned as needed?

Adam Woolfe: Oh, I think I’m a bit too old to answer this question ’cause I’ve been doing bioinformatics for too long, and I think what someone does today might not be… Or what someone might think is important today is maybe what not what I think is important. But especially with AI w- before AI came along, you had to, you, you had to do all the hard work yourself to find the answer to something. You had to look at what tools were available, what databases were available. Now you can just– AI can answer that question for you. But whether AI gives you the right question is still something that bioinformaticians– Like, having the sense to understand what is plausible and what is incorrect is really where the bioinformatics will come in the next few years, I think, because AI can give you a lot of plausible code or whatever.

Adam Woolfe: But actually, when you know what the real answer is and how the tools that it’s using should really be used, you understand that it sometimes it just chooses the default method or whatever, and that’s not necessarily the best method. So understanding the tools of the field that you’re in deeply, I think is the most important part because then you understand how that tool should be used. Because using a slightly different parameter on a tool can make the difference between having a result and not having a result. And if you didn’t know the difference and really understand how the tool worked, you would never understand why you were failing or why there was not a result at the end, or the result was not as good as it should have been.

Adam Woolfe: So obviously, the fundamentals in bioinformatics, Linux command line Unix command line Python programming knowing how a server generally works in terms of memory usage and processing and how, and all that kind of stuff. then obviously the biology, is critical understanding, like how to interpret the data you get at the end, because that is the crux. If you don’t really understand why your result looks a little bit too good to be true then you’re gonna present, that as the result and ultimately you’re gonna give people the wrong answer. So yeah, it’s a hu– The thing is, it’s such a huge field. I think it must be very, you need to learn things very quickly it can be overwhelming to know where to start. And yeah I feel for those entering the field right now because obviously it was much simpler when I started. We had like BLAST and a few CLUSTALW whatever and that was it.

Adam Woolfe: And so I’ve, I’ve spent my career learning all that stuff and that I can build on. But if you start right now, it’s just wow, where do you start? It’s huge.

Grant Belgard: How should a scientist decide among academia, a startup, or a larger company?

Adam Woolfe: I think it’s really your style. Academia is great for looking at something in a lot of depth, choosing interesting questions that don’t necessarily have a financial reward or a, a, a price tag associated with them. And I don’t think you have to choose one or the other. I think you can transition between. I think transitioning from academia to industry is very easy. Transitioning back to academia might be more difficult as, academia often relies on your your name recognition and your publication record and sometimes that’s not the priority in a biotech company that publishing is a nice to have, but it’s not a, it’s definitely not a pre- a requisite of being in a biotech company.

Adam Woolfe: If you wanna work on something that’s super interesting and but not at, at big depth, it has to be good enough, It has to be good enough to get the answer that people need to move on to to make the decisions they need to move on to the next development cycle. And if you’re okay in, in working in an environment, a high, high-paced environment where you’re not necessarily doing something that is fulfilling scientifically necessarily but fulfilling in a technical and team building way I think you can be fulfilled scientifically also biotech as well. But it’s just a very different style of working.

Grant Belgard: What’s one piece of advice you wish you had received near the beginning of your career?

Adam Woolfe: Be humble. Okay? Like we as bioinformaticians, we are at the end of the process, right? We take data often produced by people in a lab, and we analyze it, and we give the result back, right? And sometimes we don’t always appreciate the amount of time and effort that went into producing that data in the lab. The experiment may have completely failed. You may get zero reads mapping to your transcriptome or whatever. Or there’s something else, something problematic, and you just… And you can go back to the person and go, “Oh, that was rubbish. Complete rubbish.” And you don’t really appreciate, like, how demoralizing it can be sometimes to be told that that something is rubbish. And for us, it was, “Oh, we just put it through some alignment, and then it didn’t work,” right?

Adam Woolfe: And it’s like, “Okay, yeah, it’s garbage.” But like you don’t understand that person took a lot of time and effort to do an experiment and it doesn’t really help them. So communicating your result to someone is almost as important as the result itself. So ensure that you are humble and that you are empathetic and understanding and be careful with your words. It’s you can sometimes be Sometimes that can, you can come across as being harsh and and it can be problematic. So yeah, if I would say to my younger self, don’t… just be a bit easier on people.

Grant Belgard: Which unsolved question at the intersection of single cell biology, immunology, and computation are you most eager to see answered?

Adam Woolfe: Yeah. The thing that we we’re working on right now is trying to come up with models that can predict to some degree the affinity of an antibody to its antigen. This is a difficult problem. And it’s… Many people have attempted to to solve it because as I say, antibodies are slippery things. Their interactions with their antigens can be very difficult to predict. And in fact, also the place that the antibody will bind to on the antigen is also a really difficult problem. Like predicting the paratopes is not super difficult problem, but predicting the epitopes the, the amino acids on the antigen that the antibody interacts with very difficult. The success rates currently of models on that is probably something like thirty percent, fifty percent at the best. So it’s really like it’s not a solved problem at all.

Adam Woolfe: And it’s in fact a really important thing ’cause people wanna know, okay, so I have this bunch of antibodies. Sometimes an antibody can bind an antigen, but it has no effect. You wanna maybe block an active site or a receptor or something like this, and if it’s binding the site of it, yeah, it may be a good binder, but it doesn’t do what you’re really wanting to do. So knowing where the antibody is binding would be an excellent problem to solve because it would help people to just choose the best antibodies that are relevant to them. Again developability of antibodies, that’s a really difficult problem. Recent benchmarking showed that models are still really bad at predicting developability. Things like aggregation or thermostability or polyreactivity or anything like this, are still quite difficult to do and a lot of that is to do with dearth of data.

Adam Woolfe: Like antibody data is not very accessible, mostly because antibodies are so valuable, right? They sit in silos in big companies, and you don’t get access to them. So publicly available data sets that you can train off just not very available. And so models have to be based on very small numbers of of observations, which, if you’re using different technologies to measure that, it can come up with a very different result that will then screw up the model. So all these things are very difficult problems right now that I think would be great to, to solve, and we are trying to solve them ourselves. Yep.

Grant Belgard: What would you like listeners to remember from this conversation?

Adam Woolfe: I would like listeners to remember um, that life doesn’t always take the path you think it does. You may have a plan. You say, “Oh, I’m gonna be a PhD, and then I’m gonna go and gonna open up my own lab, or I’m gonna join a company,” or whatever. And sometimes, you’re thrown a curveball and things like global economy or just things that you n- would never imagine happening change everything. Even AI, I think, has thrown a lot of curveball for a lot of people, and those entry-level positions are no longer available. So don’t always imagine… Don’t, one door shuts, another door opens. Your career path will often take a very different path to the one that you imagine. And it’s, that’s not always a bad thing, and that you, it may not be something that you originally thought you wanted to do, but it can end up being something that is, that you’re like thank God that happened. I’m so happy.”

Grant Belgard: Adam, thank you so much for joining us.

Adam Woolfe: My pleasure, Grant.

The Bioinformatics CRO Podcast

Episode 89 with Beth Cimini

Beth Cimini, leader of the Cimini Lab within the Imaging Platform at the Broad Institute of MIT and Harvard, discusses her work at the intersection of microscopy, computational biology, open source software, and scientific community building.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Beth Cimini

Dr. Beth Cimini leads the Cimini Lab at the Broad Institute, where her group helps researchers turn images of cells into quantitative, reproducible biological measurements, and develops open source biological image analysis tools including CellProfiler, Piximi, and BiLayers.

Transcript of Episode 89: Beth Cimini

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to The Bioinformatics CRO Podcast. Today, I’m delighted to welcome Dr. Beth Cimini, who leads the Cimini Lab within the Imaging latform at the Broad Institute of MIT and Harvard. Beth works at the intersection of microscopy, computational biology, open-source software, and scientific community building. Her group helps researchers turn images of cells into quantitative reproducible biological measurements and develops open-source bioimage analysis tools, including CellProfiler, Piximi, and BiLayers. Beth’s career has spanned biochemistry, molecular biology, high-content imaging, and software, with a path from Boston University to a PhD at UCSF and then to the Broad.

Grant Belgard: We’ll talk about what bioimage analysis can teach us about biology, how scientific software and communities actually get built, and what advice she has for people who want to work across experimental and computational biology. Beth, welcome to the show.

Beth Cimini: Oh, thank you so much for having me. I’m delighted to be here.

Grant Belgard: For listeners who are new to bioimage analysis, what’s the core problem you spend your time trying to solve?

Beth Cimini: Yeah, so mostly we deal with images that come from light microscopes, and it feels like light microscopy should be a solved problem because we invented it in the 1600s, right? But it turns out that we invented light microscopes about 300 years before we invented computers, and so there’s a fantastic diversity of biology that we can study under the microscope. And cells might be anywhere from a pixel to a whole field of view. There’s tons of different fluorescent antibodies and can have many channels. And when you’re a bioimage analyst, you’re presented with an image and you’re sort of told, find the interesting biology here.

Beth Cimini: And so there’s a couple different pain points, one of which just being how can you find the objects of interest in the first place, which feels like it should be simple, but some of that’s because we have very good onboard neural networks, let’s call them, that help us find the boundaries of things, because otherwise we’d run into stuff all the time. Computer neural networks are still getting as good at object detection and object recognition as our brains have been, just because we’ve been doing that for a lot less time. And then there’s all sorts of sort of subtleties and nuances around, you know, how when you find the objects, can you measure them and interpret what interesting biology is happening. So it’s sort of a million dimensional problem, but it’s what makes bioimage analysis a lot of fun and very rewarding.

Grant Belgard: What does a typical week look like for you?

Beth Cimini: Oh, gosh. Well, so now I’m a PI, so now a typical week looks like a lot of emails and grants. But our lab is really fun. We’re sort of a four-legged table in that we do a mixture of open source software creation for image analysis, so things like CellProfiler and Piximi and BiLayers that you mentioned. We do image analysis methods research. So how can we make some of these things better and faster? We do image analysis almost as a CRO. So groups come to us, we collaborate with them from academia, pharma, biotech, wherever. And we do image analysis education and outreach. So a typical week, we’re touching all of those bits a little bit. And so we have an online educational platform that we’re building.

Beth Cimini: So it might be sort of checking out lessons that we’re building there, designing what the lessons are that we need to sort of put to make it easier for people to learn bioimage analysis, might be sort of checking in on a collaboration and seeing, you know, how’s it going for the workflow for, you know, these cool 3D images that, you know, our collaborator just sent over, and then sort of fixing software bugs and a little bit of everything. It’s chaotic, but it means it’s never boring.

Grant Belgard: What kinds of scientific questions tend to bring people to bioimage analysis?

Beth Cimini: Yeah, light microscopy is still one of the best things and certainly the most cost effective thing we have for doing anything where you want to do single cells. We’ve been looking at single cells under microscopes, again, for hundreds of years. So and especially anything where you want to do biology over time, we really can only of the omics that are out there only sort of reliably do sort of omics over time in light microscopy, because it’s the only omic we can do while the cells are still alive and still happy and still mostly doing their normal biology. But really, you know, we’ve estimated from some very back of the envelope things that we think probably about a third of biologists do some sort of light microscopy. So it can really be almost anything. Because we’re at the Broad and the Broad’s mission is to do things at scale.

Beth Cimini: We tend to work with people who are doing like early stage drug discovery, high content screens, but that’s a function of where we are. Biomage analysis definitely happens everywhere. It makes microscopy images hard to turn into trustworthy measurements. I think the thing that is tends to be the sort of most surprising to folks who don’t spend a lot of time thinking about it is, you know, for a lot of other modalities, we can very cleanly separate, you know, when we’re doing a sequencing read, and we see, you know, that we have a change, we have a nucleotide polymorphism that is sort of causing a biological change. Generally speaking, the change in the sequence doesn’t change our ability to do sequencing. Most biological changes won’t affect the actual measurement.

Beth Cimini: But when we get to bioimage analysis, we’re trying to look directly at the, you know, cell object or nuclear object or worm or whatever sort of object you might care about. And the phenotypic changes that are usually exactly what we want to detect, are going to affect our ability to find that object, potentially, you know, if we have a very finely tuned, you know, object detector or segmenter, to find, you know, cells, and we’ve told it, you know, cells are between 10 and 15 microns in diameter, if we have cool biology that causes the cells to be 17, our algorithm might just stop detecting the cell altogether. And so we will miss that cool biology, because it has gone outside the, you know, the definition of what a cell is.

Beth Cimini: And so it’s very hard to cleanly separate what’s a quality issue, and what’s an actual interesting biology, because we’re doing sort of detection of the things we care about, and sort of phenotype of the things we care about all in the same step. It’s not that we’re like, doing sequencing where we can measure if we’re getting a good read, and then sort of making inferences, you know, once we’ve collected our read, it’s the very things that we’re trying to detect as our biological, initial biological measurement, or the biology, there’s no intermediate steps. And so how can we do this well? And how can we make sure that we’re finding all the variability we want, and not other things like a, you know, 18 micron piece of crud that’s sitting there can be really challenging.

Grant Belgard: On that note, what makes a collaboration between a biologist and an image analyst successful?

Beth Cimini: That’s a great question. It’s one we spent a lot of time thinking about. It really, I think, just comes down to communication. We run office hours just for helping folks sort of, you know, get over particular small problems and things like that. But there’s a lot of nuances around, depending on exactly what kind of measurement you want to make. So example, if you’re taking measurements of colocalization, which are one of the most popular things to want to measure, you know, are two molecules in the same place, you know, maybe that means they’re interacting, maybe it doesn’t. There are certain kinds of controls that are really critical to do and will change the kinds of measurements that you can sort of reliably interpret if you’ve made them or not.

Beth Cimini: Which kinds of measurements are appropriate for which kinds of problems is sort of a thing that is like a passed down legacy from sort of bioimage analyst to bioimage analyst. But it’s not really anything where there’s hard and fast rules. There’s a lot of, you know, try this and, you know, under these circumstances, you know, this works and under the other circumstances. So it’s just a tremendous amount of, like, integrated knowledge that somebody who’s used to thinking about their biology as objects and not about how we turn those objects into quantitative measurements has just never contemplated. And so just lots of conversations. The good thing is all of the bioimage analysts that I know, and I know a lot of them are all, like, really, I describe us as a community of kind nerds who want to help people.

Beth Cimini: It’s really a field that tends to attract people who love collaborating and love talking about the sorts of stuff and sharing this knowledge and helping put people on the right path. It’s honestly the, like, the friendliest community I’ve ever been a part of. It’s really lovely.

Grant Belgard: How do you explain the difference between making beautiful images and extracting useful measurements?

Beth Cimini: I mean, they can be the same thing. They certainly can be. You know, a clean image that is mostly full of debris and stuff, you know, can’t, is probably going to be beautiful. Of course, beauty is in the eye of the beholder. But as a human, a lot of our, you know, visual system, we’re drawn to contrast. And so sort of people will do things like play with the gamma, which is, I’m doing hand gestures that, of course, the audience can’t see right now. Essentially, how much we change the value of the pixel that you see based on sort of the brightness of the photons that are hitting the detector, you know, because for human eyes, we like to see contrast, you might make it so that really small changes in pixel intensity sort of look very different and very striking and beautiful in an image.

Beth Cimini: But of course, we don’t want to artificially inflate the differences between two parts of our image when we’re trying to make detailed quantitative measurements. And so the things that will help you make quantitative measurements will also help make your images prettier in that, you know, they’re in focus, they’re clean of debris. But there’s a lot more we’re allowed to do if we’re just going for a pretty picture.

Grant Belgard: What kinds of projects are most fun?

Beth Cimini: Oh, gosh. I really like personally working with folks who are just getting into this space. So, you know, maybe this is the first time they’re sitting down and doing bioimage analysis, because just the like, the moment when it clicks and you see somebody be like, oh, I get it. I know what to do now. Like, that’s such a fun moment for me as a professional to sort of get to help somebody else get onto the path. I really enjoy that. But we’ve also been part of some some huge, you know, collaborative things like the JUMP-Cell Painting Consortium, which was organized by Anne Carpenter, also here at the Broad, where we were doing analysis for data made at 14 different sites.

Beth Cimini: And that was crazy to try to coordinate everything, but was also really cool to see, you know, it came out with sort of 200 terabytes of images and, you know, the biggest publicly available open source cell painting data set. And so things like that can be an awful lot of fun, too.

Grant Belgard: What kinds of projects are most dangerous to kind of, you know, be analyzed badly.

Beth Cimini: Oh, geez. That’s a good question. I think it’s really easy to do uncareful analysis, like, there’s a lot of big problems with bioimage analysis. There’s lots of hidden mine fields. I would say colocalization, I mentioned earlier, is one of the hardest ones. And that’s because there can be really subtle things with stuff like bleed through with stuff like, if I want to measure how often a little thing is inside a big thing, I need to take into account what fraction of the image is covered by little things and big things. So, you know, I have large spots and small spots. I want to know if the small spots are in the big spots, but I need to know if the whole image is big spots. They might be, but it might be by accident. It might be coincidental. And so colocalization is one of the things people most often want to do, but it’s one of the ones with the most traps in it.

Beth Cimini: So actually for this online educational resource that we’re doing, we started with saying we’re going to do a chapter about colocalization. And now we’re doing a whole course on colocalization because we know it’s really hard and we know there’s a lot of sort of subtleties to get right.

Grant Belgard: What makes cell painting useful for learning about cell state?

Beth Cimini: Yeah. So I can explain cell painting a little bit first for folks who’ve never come across it before. Cell painting is an assay developed here at the Broad in the early 2010s where the idea was, and it was, you know, more of a controversial idea at the time than it certainly is now. We have these cell images. They’re beautiful. They’re full of, you know, particular stains and particular markers to learn particular biology. But we certainly now understand from things like single cell sequencing and other just sort of computational domains, you know, just having high dimensional information, just lots and lots and lots of data means even if we don’t understand any particular specific data point, you know, we can draw connections between sort of samples or things like that. And so this is the way a lot of deep learning works.

Beth Cimini: We don’t really understand what a particular neuron in a particular deep learning architecture does, but we know when the whole thing comes together, you know, something interesting pops out. And so the idea was let’s put as many dyes as we can that will be inexpensive and fit on a standard microscope so that we can create really rich feature descriptions of cells. We’re not going to worry too much about what any individual feature means. And in fact, now in the era of deep learning, we sometimes use deep learning features that nobody has any idea what they mean. And just by doing lots and lots and lots of measurements on lots and lots and lots of cells will create big data big enough such that sort of interesting biology will fall out.

Beth Cimini: Now, I have to say, when I first came to the board, like I was trained as a sort of very standard cell molecular biologist and I was told about this assay and they were working on like the first one of the first big results papers from it. I was very skeptical. I kind of was like, there’s no way that works. Like, why did I just spend the last, you know, eight years of my PhD, like putting fluorescent proteins on particular things if all I had to do was dump a bunch of dyes on there and just sort of like study what pops out. But it turns out that, you know, making these dense feature representations is incredibly powerful. And so in that the eLife 2017 Rohban et al.

Beth Cimini: paper, which was one of the first big results papers from cell painting, sort of showing that you could detect novel pathway-pathway interactions between two different sort of gene pathways that were never previously hypothesized to interact just because they had a connection in cell painting feature space. I was like, whoa. And so now cell painting is a super well-adopted tech assay in the sort of early drug discovery space. It’s fun because it’s really sort of inexpensive to do. At scale, it’s less than $2 a sample, you know, compared to many thousands for a lot of other omics assays. It can be scaled up really easily. It’s amenable to every kind of mammalian cell we’ve ever thrown at it. And I know some folks who are working in non-mammalian cells who are using variants on it too. And you can get a lot of interesting biology out there. Now, it doesn’t work for every phenotype.

Beth Cimini: You can’t detect every kind of phenotype or every type of biology in it. But if you’re lucky enough that you can detect it in cell painting, then you have this sort of really great, really powerful assay that’s super easy to spin up and super easy to scale up if you decide you want to test it on hundreds or thousands of sort of drugs or genetic variants.

Grant Belgard: What kinds of biological variation are images especially good at capturing?

Beth Cimini: Yeah, I think there’s a lot that images are good at capturing. But, you know, a lot of stuff, because it has been the easiest modality to scale first, you know, a lot of what we know about biology comes from sequencing. And I work at The Broad, which is a place that grew out of the Human Genome Project. Sequencing is incredibly powerful and incredibly useful. But we know that there are things like, you know, regulation of translation. There’s an interesting paper that came out just a couple of weeks ago, Nature or Science, I forget which, about alternate translation. Alternate translation is actually a super common thing. Cell behaviors that happen on a fast state that happen faster than transcription or translation. You know, you phosphorylate a molecule, maybe it changes its behavior, maybe it goes somewhere else, maybe it gets sort of kinase gets turned on or off.

Beth Cimini: You can detect that with microscopy, you know, within seconds of it happening. For other omics, you have to wait for sort of like whatever the downstream activity is to sort of build up whether that’s, you know, you’re doing proteomics, and you’re sort of detecting, you know, new additions of ubiquitin or things like that. Those things take time. But for microscopy, you can do things sort of instantaneously, and you can capture cell-cell interactions, and you can capture cell behaviors in ways that are still really hard with other omics. The other omics are absolutely catching up, and all of the omics sort of working together is really our sort of vision of eventually where biology gets to.

Beth Cimini: But we are getting phenotype with microscopy at the level of like, okay, we can see how all the RNAs, all the proteins, all of the other things we mentioned in the other omics come together and make a cell behave.

Grant Belgard: What kinds of biological variation are images bad at capturing?

Beth Cimini: Yeah. So changes in sequence, obviously, are possibly detectable if they cause a downstream phenotype, but maybe not. I mean, it has to, for any sort of perturbation, we need to either have it change the cell enough that we can see a difference in something like a bright field image or a cell painting stain where we don’t know what specific changes we’re looking for, or we need to have a readout or a, you know, sort of something that’s designed to detect a sensor for a particular change. The good news is there are, you know, tons of sensors, and there’s a million antibodies in the RRID catalog, literally, but we need to know which one we want to detect. And unless you’re doing something like imaging mass cytometry or spectral microscopy, you can usually only detect four or five molecules at a time.

Beth Cimini: So if you want to see how 50 things are changing at once, that’s incredibly easy with sequencing. That’s incredibly difficult with microscopy. So parallelization is still, I would sort of sum up where I got to with that.

Grant Belgard: Where can image analysis create false confidence?

Beth Cimini: Oh, there’s a great paper on this recently from folks at Janelia called Believing is Seeing. So people assume because the computer did it, it is therefore unbiased. And having a computer analyze your images is definitely more unbiased than you sort of like doing what I was trained to do when I was sort of still a very young scientist, which was like, just like group all the pictures by like, by what they’re supposed to be pictures of, and then like decide what you think is the difference and then pick out some representative pictures and say representative image shown. That is the least, that is the least sort of unbiased you can be because a human is making those decisions. But again, things like I said about where you are making decisions about what the algorithms find.

Beth Cimini: It can be also really pernicious things like, well, I expected how, when do I decide my workflow isn’t working and decide to change it? Do I always do it no matter what? Or like, oh, in sort of run one of this assay and run two of this assay, I detected a threefold difference. And then the third time I run it, now I don’t see that difference anymore. Now maybe I decide to check, well, maybe my image analysis isn’t working anymore, but maybe it wasn’t working between one and two. And we just never checked it because the results were the same. And we were like, oh, it’s right. And so we only, you know, inspect really carefully when our results are unexpected. And so I definitely recommend reading that paper, seeing, believing is seeing, because if you believe it, you might find it. It’s a good way of how even with quantitative image analysis, you can trick yourself.

Beth Cimini: So it’s absolutely better to do quantification, but we’re humans, we’re biased. And, you know, we just have to, we have to design that into the systems that we build. And that’s something that we do try and design into the systems that we build to sort of make it easier to tell when the analysis is going wrong, but not in a way where we know the results and we expect what the results will be, but just in a more unbiased fashion.

Grant Belgard: What do you think about batch effects and imaging experiments?

Beth Cimini: Oh, batch effects and imaging experiments are like the thing that we are fighting against the most often, you know, especially with cell painting. The wonderful thing about cell painting is it is exquisitely sensitive. The horrible thing about cell painting is it is exquisitely sensitive. We’ve, you know, talking to many folks who are experts in this technique from all over the world over the years, you can tell with cell painting, which plates were in the back of the incubator versus the front of the incubator, which ones are on the incubator shelf versus like are sitting on top of another plate. And so you have all of these really exquisite differences that don’t matter, that are not what you want to detect, but they’re all mixed in with the differences you do want to detect, which are, again, the sort of things in biology.

Beth Cimini: And so being really careful with experimental design when you’re doing an imaging assay, especially something like cell painting, where you’re not always putting your controls in the same places or, you know, you’re making sure maybe you don’t even know what’s plated in each well, like you ask your friend to plate it for you and write it down. And you only sort of like you treat everything in a completely blinded fashion. It’s still really hard. And that’s why we do sometimes need to change image analysis algorithms sort of run to run or batch to batch, because there are batch differences.

Beth Cimini: And that’s, again, where that idea of like, well, is it am I now biasing the results every time that I change it, but I need to change it or it might be wrong next time, you know, you’re in this sort of like catch 22 of like, I want to keep it the same because I want my results to be comparable across batches. But also batches are different. So how can I sort of change it the exact right amount?

Grant Belgard: How do you approach rare phenotypes, subtle phenotypes, or like single cell heterogeneity?

Beth Cimini: Yeah, those are really hard. And those are things that largely as a field, we still haven’t, you know, come up with really good ways to deal with. I mean, one thing is always just kind of like, look at your data, look at your results, what level of results you can look at, you know, if you have millions upon millions of data points, you can’t look at them all one at a time, you can’t look at every image in something that has, you know, a million images, at least not without being there for three weeks. But figuring out how to sort of make sure things are, how you’re detecting the sort of rare things and how you’re making sure that you can even find them, that you’re not losing them in the segmentation process. Like I mentioned, you know, usually image analysis starts with, let me find objects that I care about, and your object detector might fail when weird biology is happening.

Beth Cimini: Hopefully, you know, something that sort of increases the likelihood of the weird biology, and you can go to look for it. But absolutely, I think we’re still only scratching the surface in terms of heterogeneity, in terms of how we deal with it, in the way that folks in like, the single cell sequencing field do in sort of large scale omics. At a smaller scale, certainly doing things like using super plots, where you can see all of the plots of your data, as opposed to like, not just like making a bar graph, bar graphs are the enemy. And sort of summarizations are the enemy can help with things like that. But it’s, it’s still something where I think we as a field need better methods. And I’m sure that, you know, smart people are working hard, and we will get better. But there’s still a lot of work to do there.

Grant Belgard: How can a team tell whether a model has learned biology rather than artifacts?

Beth Cimini: That is a great question. I wish I could answer for you. I mean, if there’s a particular piece of biology you want to, that you can induce that you know, you want to measure, then things are sort of like relatively straightforward, right? Because you sort of say, all right, I know that I should be able to induce, you know, a change in this marker, I have an assay that sort of like, you know, I have an antibody for that marker. And when I sort of add the drug or use the mutant that sort of where that marker should go up, it does. And again, doing things where you’re making sure that you’re careful to the person who designs the analysis is maybe blinded, you know, you’re not making sure that you grew all of your controls on, you know, last week, and you do all of your treatments this week, you know, basic experimental design stuff.

Beth Cimini: But especially when it comes to deep learning models, and especially when it comes to deep learning models, where we don’t have tons of data, you know, deep learning is going to learn exactly what it wants to learn. And your ability to tell what it learned can be really hard to figure out. There was a really great paper a couple years ago from somebody who had done a lot of work on cell painting stuff, and then use a tool called GradCam, they were sort of trying to classify cells with deep learning, and they use a tool called GradCam, which when you’re classifying a picture, it will sort of highlight in bright colors, the parts of the picture that it’s using to classify whatever you’re classifying. And it was saying the background is super important, not actually the cell.

Beth Cimini: And so in that case, they were able to check, they were able to sort of like, do this colored annotation of like important parts of the image. But when we teach, we have a bioimage analysis boot camp, we run a couple times a year. And we spend a couple of hours talking about how machine learning, you know, in general, and deep learning in particular, will lie to you, will learn what they want to learn, not what you want them to learn. And how can we out lazy the assay to try to figure out exactly what the model has learned? One of my favorite sort of like points to do is say, you know, what is the easiest way to train a classifier that you show it a picture of a cell and you ask it if the cell is interphase or mitotic? Just say interphase, because about 95% of cells, it turns out, are an interphase, and you’ll be 95% right. And most of us would take 95%, like 95% is a good day.

Beth Cimini: So if you just say interphase, and you don’t even interact with the thing at all, you’ll be right 95% of the time. And so trying to come up with counterfactuals, trying to sort of figure out what are ways the model could cheat, and design that into whatever algorithm you’re building, especially anything based off of machine learning are really critical, because it will cheat if you allow it to if it will find a way to do it. And we’re starting to see this now, I know, in some papers where people have taken LLM models, where it’s like, you give it an image, and it sort of tells you, like, this is a radiology image of such and such, and we see this condition in it. It turns out, if you don’t give them the image, they still say the same things. Like you you trick it into thinking it has the image, but it doesn’t, it will say the same exact things, whether it has the picture or not.

Beth Cimini: And so we need to be really careful to sort of make sure that we’re learning what we think we are.

Grant Belgard: Switching gears a bit, how do you balance tool building, tool maintenance, training, collaboration, and research? Beth Cimini Oh, gosh, I’ll tell you when I figure it out. I will say I have two fantastic team leads in my lab. Nodar Gogoberidze is our head of software engineering, and Erin Weisbart is our head of image analysis and training. And without them, you know, my life would be so much harder, and they sort of manage their respective domains really well. But yeah, there’s, and we have super smart, super hardworking people on the team, but it’s always a balance. You know, you want to be coming out with new features, but you also want to make sure that things don’t break. One of our software projects, BiLayers, is about sort of trying to promote containerization in software, because software containers are way less likely to break and are a way more portable artifact.

Grant Belgard: And so people might have to spend less time, you know, doing building, doing maintenance, because, you know, the things that they’ve made can be maintained more easily and for longer. It’s really a challenge to know, like, where in all of those things, like, am I spending my day doing the most good and helping the most people? So, so far, it’s just try a little bit of everything. But if you don’t build it, they won’t come. But if you do build it, and they don’t come, why did you build it? So you kind of need to be pushing on all of the pieces of the thing. And that’s why I like to describe it as a table. If we don’t have all of the legs of our table sort of in balance, the table isn’t a good table anymore. What makes scientific software sustainable? Oh, gosh, it largely isn’t. I will say scientific software sustainability is really hard.

Grant Belgard: And that’s a lot of it is because the vast majority of funding mechanisms are not aligned towards keeping things working. And I understand it. Like, if you’re like, well, would you like a shiny new thing? Or would you like to keep the thing that you have already working? You know, people want a shiny new thing. I totally get it. But at the same time, you know, you talk about a software project like the ImageJ project, which friends of ours make, which has about a million users a year and runs off of like a couple of people. And, you know, they’re always looking for ways to sort of fund that team better and be able to do more with that team. And this is something that a million people a year use. But nobody wants to fund something old when they could fund something new.

Grant Belgard: So I think that’s something that we as a sort of scientific community need to work on treating software as infrastructure. And we could fund infrastructure better in this country also and in the world also. But treating software as infrastructure and sort of saying in the same way that we know we need to maintain our roads and our pipes and our things in the digital age, software’s infrastructure too. And there should be more set asides to sort of keep stuff that way. You know, I mentioned containers earlier. That’s another way to at least make sure that things sort of can be used for longer. But new operating systems come out, new versions of languages come out.

Grant Belgard: And if existing tools don’t have funds, which funds equal time, then you end up with things like the heart bleed bug where, you know, open the SSL was being like done by one guy as a volunteer and something that half the internet relied on, you know, was insecure and nobody knew. So I think that’s something that we as a society could do better. And a couple of funding agencies are now starting to work in this space. But it would be great to have, you know, more focus on that being critical.

Grant Belgard: How can institutions build better career paths for bioimage analysts?

Beth Cimini: Yeah, I think first of all, they need to build career paths, not just better ones. Although I will say this seems like an area where the tide is turning. One of the first sort of big groups for bioimage analysis was the Network of European Union Bioimage Analysts or NEUBIAS, which was started in, I believe, 2015. And that was sort of one of the first times that people were like, hey, there is this career bioimage analysis. It exists. There are a few of us, like, let’s get together. And a few turned into a few hundred. There’s now a global bioimage analyst society, GloBIAS, that is trying to connect people from all over the world. But it’s a really strange career path in that, you know, it’s sort of, it’s by its definition, multidisciplinary.

Beth Cimini: You have people who started in biology or computer science or physics or math or, you know, something like that, and then pick up some of those other pieces to come to this, like, interdisciplinary space. And so it’s hard to sort of make a career path for it. But saying, hey, we have tons of our biologists making biology data, you should make sure it’s really quantitative and useful, is a thing that it now seems like more and more core facilities that I’m aware of are starting to have bioimage analysts on staff, which is, you know, a great step one. We ourselves have been running for a few years now a training program in bioimage analysis for postdocs, where we take awesome folks from the wet lab and train them in bioimage analysis. But I think definitely the sort of demand for help with quantifying images currently outstrips the supply.

Beth Cimini: But I think, you know, encouraging multidisciplinary rules, I think encouraging quantitative thinking in biology degrees, which having looked at lots of biology curricula from all over the country, you know, many biology PhDs still don’t require any statistics or computer science. And I’m not saying you need to be able to be a sort of like elite programmer when you leave, but being able to sort of have quantitative thinking be a really critical part of biology. I think there are a lot of fantastic biologists who got into biologists because it was the least math heavy science. And there’s fantastic biologists, and they do fantastic biology.

Beth Cimini: But I think as we go more and more into quantitative biology, saying, even if you’re not going to ever write a line of code, how can you structure your experiment to make it the most quantitatively interpretable is something a lot of people are never formally trained in, and they sort of pick up along the way. So I would love to see more biology programs sort of doing things like that.

Grant Belgard: What kinds of questions do people ask repeatedly? And what do those repetitions teach you?

Beth Cimini: Yeah, I think one of the most common questions that we get, so there’s an online image analysis help forum called image.sc, which if you ever, like, if you know nothing else about bioimage analysis, if you remember nothing else that I say today, remember image.sc. It is the central, like, help forum where you have more than 60 open source tools have one central help forum. And it’s a really friendly place to go and say, like, I have a picture, I want to know this about it, like, please, can somebody help me? And probably 10 people will. But we spend a lot of time, you know, sort of when we’re writing grants and things like that, chasing the hard problems, like the things that are currently really difficult to solve, you know, how can we do huge light sheet microscopy data that’s terabytes in size?

Beth Cimini: And one of the most common questions we get is just like, how with sort of really small data, can I determine what fraction of cells are positive for a particular marker, which is one of the easiest sort of image analyses to do, but people just sort of don’t even know where to start. And so I think we, it teaches me anyway, that like, we need to push the envelope on like the methods to solve things that are currently unsolvable, but that this education component of like, making sure that people know how to do the stuff that is computationally solved, but not necessarily solved for the person who needs to solve it and making it so that the tools are easy enough to use and the education is out there that people know the tools exist and know how to use them. We haven’t finished that part of the work.

Beth Cimini: And we can’t only focus on making the hot new tools because we need to make sure that the stuff that people are doing in their day-to-day lives, which is not necessarily the super hard stuff, is doable.

Grant Belgard: To talk about you, how did your own scientific interests change over time?

Beth Cimini: Yeah. For me, it really sort of began and ended with microscopy. When I was an undergraduate, sort of looking for a research lab, I was like, pretty sure I liked biology and wanted to do some research. I talked to a couple PIs and my undergrad PI, Bill Eldred, when I was interviewing with him, like pulled out a picture that they had taken of, I think this was a turtle retina, but like, you can just imagine this sort of like, brightly colored, lots of little specks all over the place, gorgeous image. And I was like, oh, that, that’s what I want to do. Like, I want to make those. And it started just as like, I wanted to do the microscopy and I still, I miss doing microscopy more than I miss anything else from the wet lab. But I got to graduate school and I wanted to answer a really fiddly, you know, thing. I studied telomeres, which are the caps on the ends of your chromosomes.

Beth Cimini: You have 46 chromosomes in each of your cells. And so you have 92 telomeres in each of your cells. And I wanted to say, you know, depending on how long the telomere is, so I have to measure that with one color of microscopy. There’s a protein that comes in two splice forms, two flavors, you know, does the length of the telomere change how much of protein flavor A versus B is there? So I have to measure three different colors very quantitatively in very small spots, a hundred of them per cell, hundreds of cells many times. And counting it by hand all of a sudden was not going to work anymore. So I had to learn to code and I really didn’t want to.

Beth Cimini: But once I did, I found the sort of like puzzle solving aspect of that to be so satisfying that I love doing bioimage analysis, I sort of picked up to the despair of my PhD advisor, all sorts of little side projects, like helping my friends analyze their data. And that when I found out when I was close to graduating my PhD, like this is a job, like I’m like, this is a job, like this thing that I love to do for fun. So I realized that collaborating with other people and helping sort of solve these image analysis puzzles, like was a huge joy to me. And that brought me to the Broad. And now I never dreamed of writing software, though, like our software that other people would use, I’d written some own code, you know, for some own software myself, I realized how much I personally enjoy and how much I personally get out of like making stuff that makes other people’s lives easier.

Beth Cimini: Like when I was a grad student, nobody was going to care at the end of the day, like what my project turned out, probably. But if I make tools that like make other people’s lives easier than every day, I’m excited to come to work, because I know that the work that I’m doing is going to lead to somebody else being able to make a cool discovery they wouldn’t have been able to before. And that for me is a way better reason to get out of bed. And a way better reason to come to work and work really hard.

Grant Belgard: What’s changed as your work has shifted from individual projects to leading people in programs?

Beth Cimini: I mean, it’s definitely now more of a, you know, you have to be thinking not just about how am I going to get through the next two or three weeks, but I have to think how am I going to get through the next two to three years and sort of really be planning for the long term because grant cycles are, you know, about a year long, you have to apply about a year before any money comes in, and you have a team of people and you want to make sure that you can keep the great people and pay them. You know, we can’t pay, especially our sort of computational folks, what they truly deserve, but we pay them as best we can with NIH budgets being what they are. And so really just sort of having this sort of like multi-level contingency plan of like, well, if I get this grant, then we’ll do this. But if we get this other grant, we’ll do that. And yeah, it becomes much more than just about you.

Beth Cimini: It’s about your whole team and how can you make sure that every single one of them has what they need. And so you can’t just be thinking about what we need now. We need to be thinking about what we need to do now to make sure we’re good a year from now, which is higher stress than just sort of thinking about what I need to do in the next three weeks. But the folks I get to work with are the best in the world. And I love getting to feel like I’m making their lives easier.

Grant Belgard: In your current role, which skills have you found matter more than you expected?

Beth Cimini: Multitasking for sure. And like fast task switching. The other thing, and it’s not a thing that I’m naturally very good at, is just documentation. Like leaving things in a place where if you need to task switch, you need to come back to something two or three months from now, like you actually remember where you were and what you did. And like, I was never the person who like loved writing in their lab notebook and sort of carefully documenting everything. But it became, you know, do this or fail, which was sort of how I came to coding too. So I guess I’ve had a lot of skills that it’s been like, well, you’re going to get better at this or you’re going to be bad at your job. So documentation was one that I really had to get much better at.

Beth Cimini: And I’m lucky to work with, I mentioned Erin Weisbart at my team already, but also Anne Carpenter are two of the best, like organized people who make the best documentation that I know. And they’ve taught me a lot as I’ve worked in this job.

Grant Belgard: Switching now to, to advice. What advice do you wish you had heard earlier? Or at least heeded earlier?

Beth Cimini: Yeah. I am grateful for all of the sort of like side projects that I got to do during my PhD. I wish I had realized earlier that the fact that I only enjoyed my side projects and not my main project meant that like a biology research career where I study a particular biological problem was not probably where my brain was happiest. And really not feeling like because that is the normal path, that is the path you must take. You know, I had very definitively decided I wasn’t going to be a PI and I came to the Broad to be a staff scientist. And then one thing led to another and ended up being a PI of a very different kind of lab, a lab where we do computational stuff. But I was, I had decided that these were the normal paths and therefore they’re the only paths available to me. And now I’ve been on this very unusual, strange own path.

Beth Cimini: And I’m happier than I think with any of the sort of normal paths, but it can be hard to imagine something besides what you’ve seen. So if you hate the things that you’ve seen, like go look for more stuff. There are weirder ways to get to a place where you’re happy than, than you possibly ever knew. You just have to find the people who’ve been on those strange paths.

Grant Belgard: I think I know the answer to this, but, you know, we always like to echo such things, right? So at what point in an imaging project should someone seek image analyst input?

Beth Cimini: Yes, early, as early as possible, preferably before you’ve done, once you’ve done your first pilot, and you should definitely be doing pilots, pilots with controls. We have a graphic in one of our papers. It’s a PLOS biology paper from 2023. The first authors are Senft and Diaz-Rohrer, where we show the circle of bioimaging and bioimage analysis. And they should be a circle. It should be that one thing feeds into the next, feeds into the next. And that your data, your image analysis and your data analysis helps you design the next best experiment. And that when you do your pilots, you analyze them all the way through to make sure that you can actually detect the statistical measurement you want to make. Otherwise, you don’t know if you can actually detect the thing you care about or not. Teams like mine have office hours.

Beth Cimini: A lot of places, your friendly imaging core will have a friendly local image analyst. If not, GloBIAS, the Global Bioimage Analyst Society, has a database of people who have agreed to be contacted on their website that you can just reach out and find a bioimage analyst near you who works on things that you work on. Because the sooner that you figure that out, the more likely that you’ll never end up sitting in a bioimage analyst’s office and then saying, the information you want just isn’t there. Those are the days I leave work the sad. It’s just when we have to tell somebody, I’m sorry, the information you want, the way the experiment was designed, we just can’t say anything.

Grant Belgard: So many parallels here with sequencing analysis.

Beth Cimini: Yeah. It’s so tricky. I hate having to do that, but sometimes the data just isn’t there. So the sooner you talk to us, the less likely that you’ll ever have that conversation.

Grant Belgard: What mistakes should people try hardest to avoid when collecting image data?

Beth Cimini: As we get more and more into people using automated microscopes, I think there will be less of this, but going in with a preconceived notion, you have to know a little bit. You have to know what stains and stuff should be present. But if you go in and you decide that you’re only going to take pictures in your controls of cells that look like this, and then you’re only going to take pictures in other fields that look like something else, you’ll find a difference. But it might not be the real underlying biology. So try to design things so that you can find more than just the preconceived idea you came in with. And it can be hard to do that while also making sure that you can find the thing that you know you care about, but it will help you sort of avoid some of those biases that will doom you before you can get started.

Beth Cimini: And that’s where I mentioned something like having your friend plate your samples for you so that you don’t actually know what’s in each well until after the experiment’s over.

Grant Belgard: How can early career scientists make invisible infrastructure work visible?

Beth Cimini: Oh, gosh, I think this is a thing that we as a sort of field still need to work on. But I mean, I think there are now more options for publishing things like the Journal of Open Source Software, like academic currency still runs on citations. And so making things citable, you know, make things like Zenodo DOIs, at least like put digital object identifiers on your work. So you can say, look at these things I made, they all have a digital object identifier. You know, some of it is, you know, administration has to meet us halfway. And you know, people who are hiring have to sort of, say, it’s not just about the papers, it’s about maybe what the papers do. But, you know, playing the game to sort of like get the metrics that you need.

Beth Cimini: Well, also, we as a community work to make it so that things like, you know, GitHub stars, or, you know, usages of tools like page visits to a blog that you wrote, that actually really helps people figure out how to do something like none of those are traditional academic currency, but all of them might be really valuable.

Grant Belgard: How should someone decide whether bioimage analysis could be a good career fit?

Beth Cimini: I think it’s really well suited to puzzle solvers. So if you love solving puzzles, I think bioimage analysis is a great career fit. I think it’s a relatively easy thing to just do, because there’s tons of images online in places like the Bioimage Archive or the Broad Bioimage Benchmark Collection. There’s tons of free tools, like, you can just sort of play and see if it feels like this is a fun thing for you. If you’re going to do it as a career, you probably have to be somebody who loves working with other people, if that’s the thing that excites you and not drains you. But I think if you like solving puzzles, and you like working with others, those to me are, and you’re organized and can sort of track having multiple things going on at once. Those for us are the things we hire for in bioimage analysts for helping them succeed.

Grant Belgard: Which bottleneck, if solved, would change bioimage analysis the fastest?

Beth Cimini: That’s a good question. It’s going to be a very, like, unexciting answer. But like, metadata standardization. I mentioned there’s like a million different things that could be in an image, and a cell could be one pixel, it could be the whole field of view. We have no common way to sort of describe what’s in an image, which means we have no common way to even say, like, should these two images look the same? Yes or no. And that goes to the thing about quality standards being hard. If we just had a common descriptive language, and smart people are working hard here. But if we all just agreed on how to describe the things, then, you know, having big databases of images and being able to make cool deep learning models and sort of suggest how your analysis should go would be so much easier.

Beth Cimini: But right now, you have to tell me and show me and we have to have a consultation and sit down and I love doing those. So I kind of don’t want to automate them away. But we could automate them away and make this a lot more approachable for everybody if we just agreed on a common way to describe things. It’s a very uncool answer, but it’s really what we need.

Grant Belgard: What worries you the most and what excites you the most about the next few years of AI in this field?

Beth Cimini: Yeah, I think I’m going to distinguish a little bit here between like deep learning in general and like generative AI, you know, large language models particularly. I think deep learning for bioimage analysis has proven incredibly powerful, especially in the context of segmentation. And people are doing fantastic work there. You know, we’ve got some papers out of like deep learning models that help you get more information about images. And all of that is great. Where I worry the most is large language models, at least the ones the common commercial ones that people use have been trained are designed to give you confident answers always. And bioimage analysis is a space that is just full of nuance. Very small changes in experimental design or in what you want to find make huge differences in workflows, make huge differences in the sort of measurements you might want to make.

Beth Cimini: And if you look at the where existing benchmarks are for bioimage analysis, and there aren’t many, LLMs are terrible at them. But because they’re advancing in other fields, because people like them in other fields, I worry about people confidently adopting tools that are not right for the job, without even realizing, doing their best, not saying like, well, this is wrong, but I’m going to do it anyway. But just saying, oh, well, the LLM says it’s confident that this is right. And I don’t think that the tools are there yet. And I think we need to understand it’s a very complex problem and not lose our skepticism. Because, you know, next probable word token predictor, you know, tells us that this is right. And those tools are powerful, people are using them to do cool things. But we need to recognize their limitations and not just sort of trust what they say.

Grant Belgard: What would you like the field to look like 10 years from now?

Beth Cimini: I want there to be slightly more professional bioimage analysts. I still think we need more of them because I think there’s still a lot of hard unsolved problems. But I do hope that a lot of the people who are currently stuck on the easy problems, which I mentioned, are actually like the majority of the problems, find it so that they can understand bioimage analysis themselves, they can learn it themselves, they can do it themselves, and then they can move on to something else. And maybe some of them will become the professional bioimage analysts who are working on the hard problems. And I think as microscopes continue to get better and better, we will come up with new hard bioimage analysis problems. But I hope that for the biologist who just wants to analyze the pictures and get on with their day, they can do that a lot more easily.

Beth Cimini: I think there’s, you know, again, we’re working on the table to make the tools better, to make the education better. And I hope over the next 10 years, I can look back and say, wow, things are a lot better than they were in 2026.

Grant Belgard: Finally, what should people remember from this conversation?

Beth Cimini: I hope what they remember is that microscopy and image analysis, like, are really powerful because they’re diverse. But therefore, it’s a discipline that’s full of a lot of like subtle pitfalls. And it’s good to talk to experts about it. But that the experts are out there, we exist, and we’re kind nerds who love talking about this stuff. So find us on image.sc, or find us in the GloBIAS Database, or come to my team’s office hours. We’re out there, we want to help you. And we really love talking about this stuff.

Grant Belgard: Thank you so much for joining us. It was a lovely conversation.

Beth Cimini: Yeah, thank you so much for having me. Have a wonderful rest of your day.

The Bioinformatics CRO Podcast

Episode 88 with Ben Langmead

Dr. Ben Langmead, professor of computer science at Johns Hopkins, director of the Langmead Lab, and founder and principal of InOrder Labs, tells us about his work building tools and resources for life scientists and his current focus on pangenomics.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Ben Langmead

Ben Langmead is a professor of computer science at Johns Hopkins University with a joint appointment in biostatistics at the Bloomberg School of Public Health. He leads the Langmead Lab, known for tools such as Bowtie, recount3, and Snaptron. Ben is also founder and principal of InOrder Labs.

Transcript of Episode 88: Ben Langmead

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. Today, we’re joined by Dr. Ben Langmead. Ben is a professor of computer science at Johns Hopkins University, where he directs the Langmead Lab with a joint appointment in biostatistics at the Bloomberg School of Public Health. His group works at the intersection of computational genomics, sequence alignment, text indexing, statistics, and high-performance computing, with a focus on open source tools and resources that help life scientists use high throughput sequencing data. Many listeners will know tools and resources associated with Ben and his group, including Bowtie, Bowtie 2, recount3, and Snaptron. He is also founder and principal of InOrder Labs and has made a large collection of teaching materials freely available online. Ben, welcome to the podcast

Ben Langmead: Thank you for having me. It’s an honor, Grant

Grant Belgard: So for listeners who know your name but not your day-to-day, what are the biggest problems taking up most of your attention right now?

Ben Langmead: We always like to work on problems where there is a hole in genomics everyday practice that can be filled with something from computer science. So in the last, six, seven years, I would say that the main thing that we’ve been working on is pangenomics. So we’re in this new era where long read sequencing is much less expensive than it used to be, the reads are higher quality than they used to be.

Ben Langmead: And so we can basically make our dreams come true in terms of assembling many new genomes. And we have our new telomere to telomere human assembly, and we have the Human Pangenome Consortium producing high quality human genome assemblies in bundles of hundreds at this point. And so this is just one of the many ways in which in genomics, we would love to move in this direction of more should be better, right?

Ben Langmead: More is better, right? But the truth is usually the more comes first and then the better maybe comes later if we really work at it. And when it comes to pangenomics, I think we’re still working at trying to make a typical everyday bioinformatics workflow truly benefit from the existence of all these very high quality reference genomes.

Ben Langmead: And to make that happen, there’s been a really superb, fun interplay between computer science and genomics, right? We needed more when it came to we need better indexing approaches, we need better algorithms, and we need it to not cost much more than it used to cost us to analyze data with respect to a single reference genome.

Ben Langmead: And computer science has been providing those missing pieces, and in fact, we now have multiple choices for where in the computer science theory literature we can turn to find new solutions to these problems. And so we’ve had a lot of fun since roughly 2019 going down multiple roads, but mainly down this road of compressed full text indexing.

Ben Langmead: How do we fit pangenomes in efficient compressed full text indexes? And then how do we let that be the basis for the next generation of tools, whatever’s coming after the current generation, but that can allow people to work seamlessly with huge pangenome collections.

Grant Belgard: How do you explain your work differently to a computer scientist, a biologist, and a clinician?

Ben Langmead: Yeah. So it’s funny because a lot of the work I do, it’s definitely more applied than what a typical theoretical computer scientist would work on. And in fact, theoretical computer scientists look at me and they see someone that’s almost like a biologist.

Ben Langmead: And which a biologist would chuckle at, right? I’m not even remotely a biologist. But to people doing very theoretical computer science, the fact that I understand the domain and understand the application of these things makes me a little bit more foreign.

Ben Langmead: And so I do, when I’m speaking to computer scientists, I tend to speak in terms of these are the open problems, and I try to frame it like it is a computational problem. The open problem is we need to be able to build a full text index, but it has to be not much bigger than the full text index for a single genome, blah, blah, blah.

Ben Langmead: And then for for people all the way on the other side in biomedicine, what I really want them to appreciate about the work we’re doing is that it should be freeing for them. It should make their lives easier. It should reduce friction. It should address the things that are bothering them about the computational work they’ve done in the past, where they say, “Oh my gosh, we’re doing all of this with respect to a single linear reference.

Ben Langmead: I know that I missed the– I know that there are entire genes that are not present in that reference. I know that there are rare alleles that are present in that reference, and yet I’m making all my arguments based on that reference. What, this is bothering me.” And so what I want the people in biomedicine to understand is that all the ways in which you feel like there’s a closed door it could be opened if only, we had an interplay and talked about what really would you prefer to do?

Ben Langmead: You mentioned at the top, our work on I’ve been talking about pangenomics, but you mentioned recount3 and Snaptron. That’s one of my favorite stories because recount3 and Snaptron, those are large scale summaries and indexes built over public RNA sequencing data sets. So these data sets, the raw data is already out there in the public.

Ben Langmead: It’s in the sequence read archive and in other resources like that. And we made summaries of that data available because it’s much easier to use than the raw data. And then we built indexes over it in the hope that we’re producing something like Google, but for archived RNA sequencing data.

Ben Langmead: And when we started publishing papers I heard from one of my colleagues at Hopkins Jonathan Ling, who wanted a meeting, and when we met, he said, “We love your paper, but it’s not what we wanted. What we wanted was this.” And so Jonathan did me the favor of making the link. He saw what we were doing, but he also saw it wasn’t what he wanted.

Ben Langmead: So he got in touch and told us what would be the missing pieces that would take our work all the way to where he was working. And so I think when the system works well there are people in each of these silos, right? The super theoretical computer science, the in-between computational genomicists like me who are maybe like glorified brokers, and then people on the biomedical research side, all understanding just enough of the other people’s language that they can phrase their problem.

Ben Langmead: And so though when I interact with people in biomedicine, I feel like my job is to try to make them understand which are the doors they thought were closed or like which are the problems they thought were too hard that if only we could solve a problem in computational genomics, they would suddenly be open and that can be more creative, right?

Ben Langmead: You can go back, ” tell me more about how you even framed your problem in the first place. Did you do, did you look at the public data right away or did you wait to look at the public data until the end to validate what you were trying to do?” So absolutely we all have different languages and we all talk to each other.

Ben Langmead: We all talk to each other in our own language, but we do especially well when we learn a little bit of each other’s language. And right now I would say those three pieces are talking to each other pretty well and I think better than when I first started in the field, which was when second generation sequencing was still young.

Grant Belgard: What has changed most about the kinds of questions that people ask of sequencing data since you started in the field?

Ben Langmead: Yeah. When I first started, second gen was really– we were really just starting to see data. So when I started grad school in 2007, and I think some of the people I was working with at University of Maryland, they had just seen for the first time a FASTQ file from Solexa or a FASTA file or something like that full of reads.

Ben Langmead: And it was just starting to sink in that they truly could not make progress with this new data type unless there were new advances in computer science. There was no hope of putting these things in BLAST. It just wasn’t gonna work. And so the problem at that point was it was almost existential.

Ben Langmead: It was how do we even make progress with this data type? But there was very fast progress. Computer scientists and others did snap to attention, and suddenly we had MAQ and we had Eland, and then there was BWA and Bowtie, and suddenly we were off to the races and the problems moved downstream, right?

Ben Langmead: So the next question was, “Oh maybe we can use sequencers to study other things. We can sequence RNAs.” And so the “what do we even do with sequencing data” sequence of questions have been pretty well answered, and we do have pretty good algorithms. We even have, what you might call like best practices and best practice workflows, and we have benchmarks, and we have comparison papers and bake-offs.

Ben Langmead: And so there’s been a lot of, I think, flourishing of improved tooling and an improved understanding of what it is you’re supposed to do when you first start with a sequencing data set. But now I’d say we’re changing our attitudes in a couple of different ways. So first of all, I think that we could change our mind about what we even use sequencing to do versus what we use sequencing data that’s already out there to do, right?

Ben Langmead: In the same way that if we start a new project today, or if not even a science project, if we’re doing a renovation to our house or we’re doing s- we’re doing something difficult today, we probably need to start to investigate things like who are the best vendors and what are the tax implications, right?

Ben Langmead: And we Google a thousand things or we ask ChatGPT a thousand things. Similarly in science, there is all this archived data, we could start there. It’s not typical, but we could, right? So in other words, even upstream of deciding what it is you want to sequence, you could decide I want to narrow my– I want to form my hypothesis or I want to narrow my hypothesis by looking out in the public data world, and we have, many petabases of of public data in the sequence, Sequence Read Archive, for example, and letting that be a a wellspring for new hypotheses or narrowing our hypotheses.

Ben Langmead: That is, I think, still a little bit of an undiscovered country, right? We, w- it’s still the case that the Sequence Read Archive is a little bit un– not, I won’t say unused, but it’s a quiet place, right? It’s not some-something that people are querying every day and learning from every day, even though we’re paying good money to have all that data out there.

Ben Langmead: Then the other way in which I think attention is shifting is that in order to build those pipelines, we made some bargains. We decided we were okay with certain trades, and I think we’re revisiting some of those now. And so all the work in pangenomics, in a sense, is really about addressing one of those maybe not so good bargains we made at the outset, which was to use a single linear reference, and that gets us very far.

Ben Langmead: The single linear reference is really, has been good to us in many ways, right? It defines a very easy to understand coordinate system. Anything you might wanna refer to is a nice, simple 1D interval. You and your friend can refer to the same interval of the genome in exactly the same language and no one’s confused.

Ben Langmead: But on the other hand, it brought us reference bias, which we’ve been living with for, a decade and a half now, where when we use a single linear reference, there are places on that single linear reference that are not a good reference for the individual we’re studying. They’re just not similar enough to the individual we’re studying, and it creates confusion from then on, and there’s often not really a way to recover from that confusion later.

Ben Langmead: There can be in some cases, but… And so I think that’s another thing that we’re turning our attention to. We built wonderful tooling, but it was built on some trades, and now that we wanna address those trades, it does require a bit of uprooting of how we did things before and a bit of disturbance so that we can adapt.

Grant Belgard: How would you explain sequence alignment to someone who uses aligners but has never thought about how they work?

Ben Langmead: Yeah that’s funny. I remember after Bowtie was out, I went to a conference, and at the conference was my colleague from Penn State, Anton Nekrutenko, who was one of the founders of the Galaxy Project. And he was saying, he basically asked me that same question because he said, “People are asking me what is the tool that is the fastq to SAM converter?

Ben Langmead: And I keep telling them it’s Bowtie, but it’s not really a converter, right? It actually has an algorithm in it.” Yeah. So I do explain this in the classroom once a year every year. And we start with the basics of, why it is that the sequencer emits fragments.

Ben Langmead: Why wouldn’t the sequencer just give us what we want, which is the complete picture from beginning to end with no mistakes? And, my analogy is it can’t do that. It’d be like you can read, but if I gave you a book and told you to read the whole book aloud from beginning to end without ever stopping, you wouldn’t be able to do it, right?

Ben Langmead: You would have a coughing fit, or you would have to go to the bathroom or something like that, right? So similarly, sequencers can’t do that, which means we’re inevitably faced with this problem of having fragments that have to be pieced together into the completed picture. And we don’t have to do this with alignment.

Ben Langmead: We can also do it with assembly, and so there’s… So I need to explain the difference between those two things, and then explaining alignment is it’s a bit like putting together a jigsaw puzzle, but where you can see a picture of the completed puzzle. And so you can hold the pieces up to the puzzle, and you can figure out how to piece it together that way.

Ben Langmead: And all of that is all well and good, but it doesn’t tell you what any of the algorithms are. So then we move into how do you do exact matching? And that’s, many decades of computer science worth of stuff we can talk about, and including some fantastic algorithms that we don’t generally use in genomics but that are used elsewhere on your computer in your day-to-day computer use.

Ben Langmead: And then we say, “Ah, but there can be differences because of sequencing errors and genetic differences.” And so then we go into approximate matching and then all our, beloved bio sequence analysis tools like Smith-Waterman and Needleman-Wunsch and such. So I guess I usually explain it through the storytelling of what are the problems we need to solve, and then why do we need this new kind of algorithm to solve these different problems?

Grant Belgard: Why does indexing matter so much in genomics?

Ben Langmead: Yeah. So indexing was in some sense the answer to the question earlier of how do we even… So like when the Solexa data was coming over the wall and people could see how voluminous it was a lot of the existing methods BLAST included, were all about scanning the reference database looking for good matches for the query, right?

Ben Langmead: You paste your query sequence into the box, and BLAST will then go and it does this in a fast and distributed way, but it’s actually literally scanning reference sequences in order to answer your question. And I think once Solexa came along, we outgrew that strategy because we had so many queries, it didn’t make sense to, for each of them, go scan the reference.

Ben Langmead: The analogy I like to use is your search engine, Google. When you put a query into Google, it is not literally scanning all those web pages while you wait and then giving you an answer because that would not be as instantaneous as you need it to be. It’s using an index, right? So indexing comes into play when you have one reference that you know you’re gonna use again and again.

Ben Langmead: I’m gonna go searching for a needle in this particular haystack over and over again. And not only that, but I’m gonna search for many needles in that haystack. When you get to that point, then it becomes a, a question of amortizing. In computer science we like to amortize effort.

Ben Langmead: So if we’re gonna put in effort ahead of time to build a data structure to help us solve this problem, it makes much more sense to build it over the haystack. And so building that data structure to help us search inside a very large collection of text that’s what we call text indexing.

Ben Langmead: And the name index is a good one because it’s a good analogy. The index of a book is a perfectly good analogy for what it is we’re doing, right? We’re building something that has pointers into the huge the vast amorphous data set that we’re trying to index. It has pointers into it that help us jump to exactly where we need to go at query time.

Ben Langmead: And so text indexing has, in computer science, has been around for a very long time. There’s a wonderful sort of parade of data structures that came out over time, and genomics has absolutely been on the bandwagon of following and adopting each of these as they became popular. So suffix trees and suffix arrays were already very popular when I started graduate school, and then through the work of others and through my work in graduate school, these more, I won’t say modern but very frequent, very frequently used today this very frequently used idea of the Burrows-Wheeler transform became the one that really was useful because it was the one that let everything fit in memory on a typical computer at least at the time.

Ben Langmead: And these days now the trick is trying to use some of those same tools to get entire pangenome indexes to fit in memory on a typical computer.

Grant Belgard: How has your thinking about alignment changed as sequencing technologies have changed?

Ben Langmead: Well, sequencing technology, the way it has changed it didn’t have to do this, but it did in fact reveal new problems to solve as it went. I should say yes and no, because sometimes it reveals new problems and sometimes old problems become current again. So for example, long-read sequencing produced reads that were more amenable to the kinds of alignment algorithms that had already been developed for Sanger sequencing.

Ben Langmead: And therefore part of the need from the computational genomics community was let’s rediscover and dust off some of those older algorithms that we had for longer reads. So sometimes it points us to something we’ve already done, but sometimes it shows us that something entirely new is required.

Ben Langmead: And so these days, in a way, the advent of long reads has changed everything we want to do, whether it’s with short reads or with long reads when it comes to alignment, because the long reads gave us this abundance of new references. And the trick is: how do we use them? How do we use them in a way that truly benefits us, but for not much more cost?

Ben Langmead: And so that is a truly new… It’s in one sense, a truly new way of viewing the problem that we have to solve. But the good news is it actually calls on some of the exact same methods that we were already using, right? So like the Burrows-Wheeler transform, which I mentioned before, this is like the core.

Ben Langmead: If you’ve used BWA or Bowtie, the Burrows-Wheeler transform is at the core of those tools. It’s at the core of some other genomics tools as well. It has many variants like the positional BWT and things like this. But the BWT was originally invented not for indexing, it was invented for compression. So it was invented to basically compress files on your computer, like Zip does.

Ben Langmead: Which is a fortunate thing because this problem of indexing pangenomes, it requires compression, right? One thing that won’t work is just building the same old indexes over each of the sequences in your pangenome. Let’s say your pangenome consists of hundreds of human haplotypes, because that’s just gonna grow proportionally to the length of all those haplotypes, and you will very quickly exhaust the memory of the computer you’re using.

Ben Langmead: So it turns out, happily, that the exact tool that we’ve been using since, 2009-ish for our, to solve alignment problems, the Burrows-Wheeler transform, was invented for something like this in the first place. Compressed representations. So the latest revolution is really about compressed indexing and being able to efficiently query compressed indexes, so where the index itself is not much bigger than a fully de-duplicated version of the text that it’s indexing

Grant Belgard: Where do pangenome approaches feel mature and where are they still experimental?

Ben Langmead: That’s a tough question to answer. So I think that there are some things that people are getting used to doing that involve using multiple references. But whether people are yet used to fully pangenome approaches, I don’t quite know, right? I think people still wonder this, right?

Ben Langmead: They still wonder “how do I do this in my day-to-day?” They know there’s such a thing as pangenomes and graph pangenomes, and there’s graph aligners, which are probably the most mature single thing we can point to. So things like VG and Giraffe, these are quite mature, so people can use these tools.

Ben Langmead: But what people do day-to-day still looks to me like thinking of genomes as individual things, but trying to use more of them. So instead of embracing the whole forest, it’s like just swinging from vine to vine or something like that. So let’s try to use this reference genome, and if it looks like the situation is breaking down in some situations, like it looks like some of my reads failed to align, but they did align to this other reference genome, that means that other reference genome is the one I need to study this particular thing.

Ben Langmead: Maybe therefore I’m gonna hop from this reference to that reference and, but, and do everything again, but with respect to that other reference. Which isn’t even a very bad way of doing things because you, it doesn’t take a large number of reference genomes before one of them is probably a pretty good reference for whatever it is you’re trying to study.

Ben Langmead: But it’s definitely not yet the case that we train people to use pangenomes instead of linear reference genomes. So that’s, I think, one thing that eventually needs to change, is it has to be part of our training. We have to learn what they are and what reference bias is, how to detect when we might be falling prey to reference bias, and then understanding how to use the pangenome to cure that.

Ben Langmead: But also there’s a lot of unanswered questions, right? So there’s actually just plain still more work to do in pangenomics. And, I think what a lot of biologists would say is, “It’s great that you have the pangenome for me to use, but how do I explain it? If I use it in my paper, how do I explain what I found?”

Ben Langmead: There’s a anxiety about the notion of pangenome coordinates. Like, how do I even describe where I am in a pangenome? And is it gonna make sense to somebody who reads my paper? Compounded with that is the fact that pangenome coordinates themselves are a little bit mythical, right?

Ben Langmead: It’s not necessarily true that a pangenome does or should have a global coordinate system because it’s not necessarily globally collinear. If you insist on understanding a pangenome in a global way, the global landmarks are gonna fall away more and more the more genomes you add, right?

Ben Langmead: Because some genome is gonna disagree about what’s next to what in the genome, and as soon as you add that one, you’ve knocked out part of the common coordinate system. So part of what we need to address is maybe we don’t need to insist on the pangenome coordinate system being truly global.

Ben Langmead: Maybe it’s something that we navigate in a way where it would be awesome if we could answer our questions in a fully zoomed out global way. But when that’s not possible, we should zoom in and say, “Okay I can provide a good coordinate system for what it is you’re studying, but only if I limit my attention from here to here and discard these sequences from the database.

Ben Langmead: Now I can give you a really clear, nice picture, maybe like a multiple alignment or something like that of that area.” But these tools are also missing, right? In fact, at this point, the pangenomes are getting to be so big it’s hard to even build a multiple alignment of them, right?

Ben Langmead: We already have pangenome data sets for which I doubt we’re ever gonna have a good, reliable multiple sequence alignment of those sequences, even though they are as related as humans, for example. So another obstacle is how do we both take advantage of the additional data that the pangenome’s giving us, but without going too far, right?

Ben Langmead: So more is better right up until the moment that it’s not, right? More is better right up until the pangenome stops being a good global coordinate system for what you’re trying to study. So we do need a way of saying, “More is better, but there are limits. And if you wanna, if you’re studying what you’re studying, we recommend this subset.

Ben Langmead: Or in this area, you should be using this subset. Or I see what you’re studying I’ve detected that these 80% of the references are perfectly good for what you’re trying to study, and these 20% are just gonna be confusing, so I’m gonna zoom in for you.” There’s still a lot of work, I think, to be done in that area.

Ben Langmead: So I don’t know if it’s even time yet to train people to learn about pangenomes. It’s certainly a good time to learn about reference bias and how you can avoid it by moving from linear reference to linear reference, and there are certainly some problems we can already solve with graph pangenomes and graph aligners.

Ben Langmead: But the future is still fairly uncharted, I would say, when it comes to how are people really going to switch to pangenome thinking, right? We’re still trying to wrap our heads around what pangenome thinking is, and therefore how we should, what tools we should build and how we should train people.

Grant Belgard: How do metagenomic data problems differ from human genome data problems computationally?

Ben Langmead: In some ways they’re quite similar. So for example, in metagenomics, maybe you’ve collected a sample from the human gut or from a soil sample or a surface that you’ve swabbed in a hospital or a waterway, and you are sequencing everything that is there. So you are getting sequencing reads derived from the DNA genomes of whatever was there.

Ben Langmead: And again your job is to put together a puzzle. So quite similar to if you had just sequenced one human. But the problem with this puzzle is you don’t necessarily have the picture of the completed puzzle for all these sequences, right? Because there’s a lot of organisms that you could collect in a metagenomic sample for which we just don’t have reference genomes and maybe don’t even have nearby.

Ben Langmead: Maybe we don’t even have cousins, we don’t have anything close enough in our reference databases. So that’s one immediate difference. Another difference is it’s quantitative in the sense that if a organism is there in a higher abundance, you’re gonna get more data from it that usually when we’re sequencing, say DNA from a human, we want to sequence equally from all parts of the genome and that’s quite possible.

Ben Langmead: But for a metagenomic sample, we usually have to deal with the fact that some organisms are much more abundant than others. That’s another additional problem But other than that, it actually looks similar, right? Because we usually analyze these metagenomic sequencing reads with respect to a reference database.

Ben Langmead: That reference database now is not just one species of reference genome, but it’s a whole collection of species of reference genomes. But again, one of our most valuable tools in this situation is indexing. So this again is a text indexing problem. It’s even, again, a compressed text indexing problem because if you think of what reference genomes we have for what clades in the tree of life, we have a huge abundance of reference genomes for some clades and then poorly sampled reference genomes for other clades.

Ben Langmead: But like things that are important to human health, things like COVID and SARS-CoV-2 and E. coli and things like this we have tons and tons of reference genomes in those clades. But then in other clades, maybe for organisms that are hard to culture or hard to isolate, we don’t have as many reference genomes.

Ben Langmead: So we have some parts of the tree of life where it is crucial that we do compressed indexing because we’re just not gonna be able to fit all those sequences in unless we do. And so indexing, compressed indexing, these are still the valid problems. We end up using, in practice, often an expedient here, which is to build a k-mer index.

Ben Langmead: So in other words, the sort of keys, if you think of the index of a book, right? The keys there are the key terms that you might wanna look up. In a k-mer index, the key terms are just substrings of length K, right? So those are the key terms that are held in our metagenomics index. So that’s, it has its pros and cons.

Ben Langmead: K- Once you do this, you have to pick the value of K, right? So you have to decide what is the true one and only substring length that’s gonna tell me what I need to know. And, almost everybody just picks a default, which is often 31. But it is far from obvious that’s what should be done.

Ben Langmead: It’s also far from obvious that’s what should be done everywhere in the tree of life, because again, some parts of the tree of life are very dense with sample genomes and some are very sparse. Some parts of the tree of life just have more divergence among the individuals in that clade, and others have less.

Ben Langmead: So we use these expedients, and that’s because it’s hard. It’s harder than, for example, the human pangenome, because the different clades don’t have sequence similarity across the clades, right? So the compressed indexing is only going to be beneficial within a clade. It’s not particularly beneficial across clades.

Ben Langmead: So we have to rely on these expedients. We can’t quite do anything quite as comprehensive as a full text index. At least not usually, though some tools do this, and I think the future probably is for tools that do this. So it is in some ways a harder computational problem, but it is a similar one

Grant Belgard: What makes a bioinformatics tool trustworthy?

Ben Langmead: So the research world is appropriately a world where we like to let 1,000 flowers bloom and a lot of software gets put out there. And this is good, right? Students need to be able to have an open playing field to go ahead and make their contribution.

Ben Langmead: But not everything is a hit, right? And it’s impossible to predict what’s gonna be a hit. It’s slightly possible to predict what’s gonna be a hit, but generally it’s not possible to predict what’s gonna be a hit. So the truth of the matter is everything starts on kind of equal footing.

Ben Langmead: Maybe this is gonna have traction, maybe it’s not. And if something has traction, then hopefully the student and the advisor are seeing that happen and responding. And often the way you respond is you say, “Oh, wow, we should be thinking about releasing new versions and talking to some of the users, or at least trying to figure out what the users are thinking.”

Ben Langmead: And sometimes things get in the way of that happening. Maybe the student graduated and the PI is very busy and teaching, and so nobody notices. This is okay. But sometimes everything comes together and we can see, oh, this is a hit. Let’s put out more indexes, for example. Maybe this is an indexing-based tool and some people wanna use human indexes, but other people wanna use mouse indexes, so let’s put out more mouse indexes.

Ben Langmead: Or maybe people are requesting features, let’s work on those features. And you can tell from the outside, to your question, you can tell from the outside whether this is happening or not, right? There’s a lot of signals, right? Has the website been updated? Is there a recent release?

Ben Langmead: Are they, are there, is there anything going on the issue tracker, right? And I say this and I think, I hope people don’t check all my issue trackers on all my tools ’cause I’m not proud of what you’ll find in all cases. But we have ways of from the outside interrogating to figure out if this is having a long and healthy life in this lab, or if maybe the lab has turned their attention elsewhere and maybe this thing isn’t too usable.

Ben Langmead: Documentation is another very important thing. Has the team done something more like a reference implementation of an algorithm and then moved on? Or did they think through with empathy for the user? What will the user actually do and be confused by, and can we break it down into subcommands that the user understands what they are and how to sequence them in order to get to their, to their goal?

Grant Belgard: Your lab has always been really good at that, by the way.

Ben Langmead: I appreciate that and it’s ongoing communication, right? And sometimes it has to be face-to-face, right? There’s definitely tools where I think everything is going a certain way. I’m interacting with the issue tracker. I’m having weekly meetings with my software engineer.

Ben Langmead: Then I go out and give a talk and have a meeting with somebody, and I realize, oh, this person thinks about what they want from this tool in a totally different way than I expected. And then, next meeting I have with my software engineer, I’ll relay this and I’ll say “maybe this means that we should provide this additional command, or maybe this means that we need to take it in this direction.”

Ben Langmead: So it’s ever-evolving and it definitely requires nice high bandwidth communication with the community, which doesn’t always mean it can be limited to things like GitHub issues or emails. It sometimes has to be go out in the community and ask people questions and figure out what they’re confused about or what they’re worried about.

Ben Langmead: But yeah, we all know the difference and my take would just be it’s okay for software to go out there and not get traction and gently, sink back into the earth over time. That’s perfectly okay. I think that is important for a healthy research ecosystem.

Ben Langmead: PhD students need to have projects, and nobody knows which are the ones that are gonna get traction. And so if everything works out just right, then the ones that should get traction do, and the people who made it have the time to pay attention and do the documentation and have the empathy to make it truly easy to use and also commit to it long term and go out and talk to people about it when they’re at conferences and talks and such.

Grant Belgard: What do you decide when software is ready for other people to use?

Ben Langmead: I think it is true that some of what we put out is more of the form of a reference implementation. So in other words, that software is there because the idea is good, and in order to show that it was good, we had to implement some software. And that’s not always software where I’m gonna expect a white coat-wearing biologist to come along and download it and use it.

Ben Langmead: And again, that’s okay. There should be a notion of a reference implementation. And in the day of AI coding and AI rewrites, it could even be that reference implementations are the main thing that people like me put out. Because if it’s gonna get traction, perhaps it’ll get traction through an AI rewrite of my reference code rather than my code per se.

Ben Langmead: I like to think… first of all, when you write the documentation, that’s a pretty good stress test of whether you really did think through how users are gonna use it, right? ‘Cause if your documentation is reiterating your paper, that’s probably not a good sign.

Ben Langmead: But if your documentation is saying, “If you have this problem, then run this command,” and you thought through what are the problems the user might have, then that’s better. So it becomes like a question of what are the nouns and the verbs? Do they map onto things that biologists actually already know?

Ben Langmead: Can we use standard file formats like when a biologist comes along to interact, are we gonna be introducing them to new file formats? That often is a point of friction. Or are we telling them, “Yeah, you can use the stuff you’re very used to using.” You can think of it as a converter if you want to.

Ben Langmead: It’s true, we put a lot of effort into the algorithm, but sometimes things really do fit into a pipeline and play a role in a pipeline, and we need to allow the user to think in terms of getting from point A to point B in their pipeline and not get too hung up on what’s our algorithm. So yeah, I think that if you can think in terms of the user’s problems and nouns and verbs, and if your documentation is written like that as opposed to, “Here’s everything we did,” ” here’s all the functions we implemented,” then I think you’re well on your way.

Grant Belgard: What first pulled you towards computer science?

Ben Langmead: Gosh, computer science I was into from roughly high school. Maybe even middle school. So I was at a middle school where they taught us Basic, the programming language, and I could tell that it was fun, and I could tell that I was good at it. And so when high school came along, I took all the computer science courses I could.

Ben Langmead: Also, my brother, who’s eight years older, he was taking computer science classes in college, and I was asking him to send me his homework so that I could work on them too, and, I was trying to learn C programming. And so by the time I got to college, I was very sold on the idea that I really like the ability to tell the computer what to do, and that it truly does what I told it to do.

Ben Langmead: And then when I got to college, I knew right away I wanted to major in computer science, and I took my computer science courses. And I think what I further fell in love with was computational thinking and thinking about software and hardware. You mentioned high-performance computing as one of my emphases at the top and that’s true.

Ben Langmead: In fact, before I went to graduate school, I basically only did high-performance computing. And my favorite thing about that is like un-un-black-boxifying the computer, right? Being able to think of the different pieces of the computer so that it’s not a mystery. “I ran this, I was sure it was gonna be fast, but it was slow.

Ben Langmead: Why did that happen?” Trying to demystify these things and say, “No, I get it. I see how that mapped in an unfortunate… It was a good idea, but it mapped in an unfortunate way onto the way the hardware works. I need to think about the way the hardware works and come up with a new idea.” So I was, I’d say, sold from an early age on the idea that programming is fun and it’s just a fantastic hobby, and I, and if I could do it as part of my job, that would be great.

Ben Langmead: And then in college, I think I fell even more in love with making the computer something other than a black box. And then the next step for me was graduate school, which is where I truly got to commune with the idea that computer science can be used for basic science, which is not necessarily something that computer scientists ever learn, right?

Ben Langmead: Like you can go through an entire computer science undergraduate curriculum and think the whole time that computer science equals big tech. Computer science equals Google, Microsoft, Facebook and of course I’m gonna move to San Francisco when I’m done, right? But going to grad school, at the time I did especially, taught me no.

Ben Langmead: People who are computer scientists can work shoulder to shoulder with basic scientists, and that was the next step in my evolution. And that piece I think is… like my story is a very common story, I’d say right up until grad school because that final piece of connecting computer science with another science is not something we do so well in our curricula.

Grant Belgard: Was there a moment when you realized that genomics had problems you wanted to spend your career on?

Ben Langmead: Totally. Yeah, when I got to graduate school, I was working with one professor at the time at University of Maryland, Bill Pugh, and I was not working on genomics. But at the orientation events, I went to an orientation event where Stephen Salzberg was speaking, and he was describing next-gen sequencing which at that time was brand new.

Ben Langmead: And like I say, he had probably only seen one or two, Solexa data files at that point. And he was talking about what a big deal it is that it’s so inexpensive now, and what a big problem it is that we can’t– we don’t think we currently have the tools to analyze these. And I was so compelled by that, that I went up afterwards and I asked him “look, my background is this and this.

Ben Langmead: Like I, I kinda do high-performance computing, and I’m very much a computer scientist, and I don’t know, last biology class I took was freshman year of high school. Is this an area where I can work?” And he said, “Absolutely.” He was like, “We need people like you.” And he told me to talk to two of his students at the time, Mike Schatz and Cole Trapnell.

Ben Langmead: And luckily for me, Mike and Cole were giving a seminar shortly after that on something they’d been working on called MUMmerGPU, which was a whole genome alignment program that used CUDA and GPUs. This was in, the year 2007 or 2008, so it was very early for that kind of thing. And when I went and saw their talk, they gave a tag team talk, I was at that point completely sold, right?

Ben Langmead: I was– I wanted to be like them. I would do anything to be like these guys. And so I started to talk more to Stephen and also to Mihai Pop, who was my master’s advisor along with Stephen Salzberg, and try to figure out, okay what are the handholds we can get on this alignment problem?

Ben Langmead: And by the first summer after I started grad school, I was fully only working in that area.

Grant Belgard: For a student who likes both biology and computing, how should they choose what to learn deeply?

Ben Langmead: That’s a great question and I can only step on toes by answering your question, so I’m just gonna go ahead and do it. So what I always tell students who approach me and ask that is, I think the computational toolbox is the one to learn earlier. If you’re gonna learn– If you’re gonna sequence these things I’m gonna do one and then I’m gonna do the other, or I’m gonna do one and I’m gonna apply it to the other.

Ben Langmead: I think the thing that’s important to get early is the computational thinking. It is possible to then pick up more of the biological thinking. Okay? That’s my opinion, right? This is not what I would say necessarily to my biology colleague, but this is what I say to students. So I feel like if you’re, if you have that computer science training and toolbox, there’s not really a door that’s closed to you.

Ben Langmead: I don’t think there’s too many professional opportunities where they’ll say, “But you’re a computer scientist. We’re, we do XYZ.” As long as you can demonstrate that you are, that you have the computer science, but you know how to apply it and you know how biological people think, I think it’s perfectly fine to have that computer science degree.

Ben Langmead: So I’m often talking to students who are trying to weigh, “Do I wanna major in computer science and then minor in something?” Or maybe they’re looking at grad school and they’re saying, “Do I wanna do a computer science PhD program or maybe a, a biomedicine or a computational biology or some other…”

Ben Langmead: I usually tell them, “I think that computational toolbox is the crucial one. It’s important to learn it early, and so it’s good to put your academic emphasis on that and grow into the other areas.” I think, like one of my experiences early in my career was before I got my PhD, I worked as a research associate at Johns Hopkins in biostatistics with Rafat Irizarry, a biostatistician, and he was collaborating with a epigenomicist, Andy Feinberg.

Ben Langmead: And I would go to Andy Feinberg’s lab meetings, and at first, I understood basically none of what I heard. It was like being immersed in a foreign language. But if you keep going, and maybe also if the students are willing to answer your questions after the lab meeting if you didn’t understand something, you start to see s- things, certain things repeating, you start to see certain themes, and then you do pick it up after a while.

Ben Langmead: And so I feel like that is a possible thing to do, whereas I’m not sure quite how possible that is with computer science. You have to bang your head against, “My program is not compiling. My program is not working. My algorithm is taking 10 times longer than I thought it would.” Those are things you have to commit to long periods of time of banging your head against it before you can get to the end.

Ben Langmead: So I usually advise that the toolbox to get early and to have represented primarily in your credentials is the computer science one. Now, I’m obviously biased and I always tell students that. So listeners should know that was a biased opinion.

Grant Belgard: To close this out what advice do you find yourself giving again and again?

Ben Langmead: Okay, so I interact with a lot of students at all levels, undergrad and graduate students. And so I think when I’m talking to undergrads, it’s usually because they’re trying to figure out how to map out their college experience and what do you do? And they’re trying to figure out, should I cram in as many classes as possible, or should I triple major?

Ben Langmead: And what I’m usually telling them is something like it’s actually quite difficult to predict what it is you did as an undergrad that’s gonna unlock that next door in your career. It might seem like you want to take as many classes as possible in the thing you’re most interested in, but my story is when I finished undergrad, I applied to several jobs, and I ended up getting a job at a small consulting company called Reservoir Labs.

Ben Langmead: And after working there for a couple years, I was talking to my boss, and he revealed, “Basically, the reason we hired you, the reason we picked your resume out of the pile, is I really liked your writing sample.” And this was, like, a essay about Japanese history from one of my East Asian studies classes, right?

Ben Langmead: So could I have predicted that my first job out of undergrad as a computer science major would have been based on my writing sample? No. And so I think that as important as it is to craft the right undergrad curriculum, it’s also important to explore and make sure that you’re developing that wider skill set: communication, writing presentation skills, as well as whatever it is you’re doing.

Ben Langmead: And then for grad students, I would say one of my, I have lots of advice for grad students, what does it mean to be a PhD student? But I think one of the messages that’s hard to take but important is failure is an option, right? This is science.

Ben Langmead: Your most beloved hypothesis the underlying science doesn’t care how much you love it, right? It’s you have to go with what the data is telling you. And unfortunately, that means that you can’t point at the calendar and say, “This thesis chapter is gonna be done by this date,” or at least not not without being frustrated at the end of the day.

Ben Langmead: So I think people who go to grad school are committing themselves to allowing science and the data to tell them when they’re done and when they’re not done, and that’s a tough commitment, but it’s the right one.

Grant Belgard: Ben, it’s been a great conversation. Thank you so much for joining us

Ben Langmead: Thank you. Pleasure to be here

The Bioinformatics CRO Podcast

Episode 87 with Elliott Margulies

Elliott Margulies, Director of Bioinformatics at BillionToOne, discusses the application of molecular diagnostics to clinical testing. 

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Elliott Margulies

Elliott Margulies is the Director of Bioinformatics at BillionToOne, a molecular diagnostics company applying quantitative approaches to prenatal screening and liquid biopsy.

Transcript of Episode 87: Elliott Margulies

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. Grant Belgard, joined… i’m delighted to welcome Elliot Margulies director of bioinformatics at Billion to One, a molecular diagnostics company applying quantitative approaches to prenatal screening and liquid biopsy. spent his career at the intersection of genome technology, computation, and clinical translation, including earlier work at NHGRI and Illumina. In this episode, we’ll talk about what he’s working on now, how his career developed, and what advice he has for scientists and engineers who want to build bioinformatics that matters in the real world. Elliot, welcome to the show.

Elliott Margulies: Thanks, Grant. It’s really great to be here. Great to see you again too

Grant Belgard: uh, The pleasure is mine. for listeners meeting you for the first time, what are you working on now and what problems are most alive for you day to day?

Elliott Margulies: Well, I’m like you said, at a molecular diagnostics company, and so day to day uh, my team focuses on the processing of uh, data uh, for prenatal testing. And we focus on everything from kind of the infrastructure uh, and analysis of kind of primary sequence data, making sure the instrumentation is running correctly, our QCs are, are appropriate, all the way through to building and maintaining curation and interpretation systems and building out reporting platforms.

Elliott Margulies: We’re, we’re kind of the glue across all the different parts of the organization that has to touch a sample from when it’s, uh, acquired to when it’s reported out, uh, whether that be the lab, our genetic counselors, lab directors, uh, even some software teams who are managing, uh, interfaces, uh, and portals that, uh, our physicians and patients interact with.

Grant Belgard: Would you explain the core bioinformatics challenge in your current work to a genomics audience that’s new to this particular area?

Elliott Margulies: Well, we very much focus on actionable, informative, places in the genome. Uh, This is actually, uh, a kind of a, a migration that I’ve made in my own career, uh, from, you know, figuring out algorithms and methods to interpret and analyze whole genomes, uh, bringing that to scale and thinking about how healthcare systems might adopt that, to now being at a company that focuses very much on very specific parts of the genome, uh where, um, we are able to determine with high confidence, uh, whether risk variants or other de-deleterious variants are present.

Elliott Margulies: Uh, and we, and, uh, we can do so both by looking at, um, the DNA from mom, um, but we can also look at the cell-free DNA, uh, that represents the growing fetus inside of mom, uh, and determine the probability that, uh, the baby is carrying zero, one or two, uh deleterious alleles.

Grant Belgard: Can you make technical work understandable across teams with very different vocabularies?

Elliott Margulies: Oftentimes we remind ourselves that everybody is in a silo to a certain extent with their own lingo, and we make it, uh, comfortable for people to pause and say, “Wait, I don’t understand you,” or, “What was that acronym you used?” And, uh, we really have a fun culture of kind of acknowledging that everybody is an expert in their own field, but when th-those interfaces start to talk to each other, you gotta slow down and, and not be afraid to say, “I don’t understand.”

Elliott Margulies: It’s actually a really strong part of our culture, um, at this company right now where that understanding of constant learning and explaining or, um, feeling comfortable saying, “I don’t understand,” or feeling comfortable saying, “Ah, I think I may have screwed that up,” uh, you know, “Let me fix this,” or, “Let me learn what, what happened with the process that we can improve upon.”

Elliott Margulies: But it all starts with, uh, clear communication and feeling comfortable, uh, at those interfaces and slowing, slowing the conversations down, up-leveling, using analogies

Grant Belgard: What do you wish more people understood about the bioinformatics behind clinical testing?

Elliott Margulies: We’re, incredibly rigorous, right? So when I think about the reproducibility, and the confidence we have to have that all of our pipelines are working, we’re constantly thinking about how can we cross and double and triple check that everything is working correctly, that we’re giving the, the right answers to the right people all of the time.

Elliott Margulies: And so, the– running a pipeline is easy, but running a pipeline at scale with a level of confidence that every patient is getting the right answer, as often as possible, and that we have ways of… ‘Cause oftentimes it’s the exceptions that we spend the most time on, right? So, fortunately in a prenatal testing situation, most samples are negative, and that’s g- that’s a good thing.

Elliott Margulies: or I guess another way of thinking about it is that most of the time, sequencing runs and analyses work out just fine. But, we spend the majority of our time on those, those small exceptions, whether it’s a complicated case, a sequencing run that didn’t go as well as we thought, some other issue that’s kind of coming into the pipeline.

Elliott Margulies: We spend the majority of our time, right, thinking about, you know, those small number of samples that are not within the, the norm.

Grant Belgard: When you encounter those, edge cases, are they typically treated, as, one-offs, or, do you typically make changes to the pipeline or process or something on the basis of those? Or are they just so rare and weird that they, typically don’t trigger kind of

Elliott Margulies: No, I’m actually gonna quote, a team member of mine, his name is James Hart. we were recently talking about, building new pipelines for future products. And, I s- I, we were talking, and I had said something to the effect of, you know, “We need to make sure we can handle the edge cases.” and he looks at me and he says, as he’s, has a great, mindset when it comes to building the- these types of, infrastructure.

Elliott Margulies: He says, “Elliot, there are no edge cases. There are only workflows that are run rarely.” And it was a, it was a nice epiphany, I thought, where, And it was, I was like, “Oh, we’ve got the, the, the right person thinking about this.” But it is true in the, in the sense that for some of the legacy systems that we manage, when there are edge cases or one-offs, we see them with some level of repetition.

Elliott Margulies: And particularly as things scale, it’s important for us to build systems that can more handle those types of situations in, in a more standard way.

Grant Belgard: So I guess on that note, how do you decide whether a problem needs a better assay, a better model, better data, a better process?

Elliott Margulies: the, you know, there are many types of rules and kind of prioritization, right? So you, you have, you know, a finite set of really smart people on your team, who are balancing both kind of production responsibilities as well as building new products, as well as automating and scaling kind of existing infrastructure.

Elliott Margulies: And we balance our time across these things, and oftentimes we l- we look for, you know, the proverbial low-hanging fruit or take that 80/20 rule. You know, where, where can I spend a small amount of my time and, and have that be a large impact? And oftentimes we think in that, in that way when we prioritize work, particularly over the short term.

Elliott Margulies: And so we try and balance the short term, where can we do a small amount of effort that’s gonna have a big gain? and a- another metaphor sometimes we talk about is those types of small efforts that have large gains start to get a flywheel going of, more automation, and that more a- additional automation frees up even more time to then automate even further.

Elliott Margulies: And before you know it, you’re starting to, to spend time thinking about the larger system and the longer term growth that, that is necessary.

Grant Belgard: Yeah. How do you think about uncertainty when a computational result may affect a clinical report?

Elliott Margulies: We’re very much a first principles-based company, and so a lot of the analysis that we do, uh, is rooted in statistics, uh, and probability and, and quantifying, um, you know, where we, we have the ability to essentially count molecules, um, at a– that are at a very low level. Uh, but we can do so very– in a very quantitative way.

Elliott Margulies: And so oftentimes we are, uh, we, we put probabilities around, um, all of our calls. And, uh, we are constantly looking at, um, outcomes data. Uh, so we have a fantastic outcomes team that, uh, of course it takes, you know, nine months or so to, uh, to get the, uh, the outcomes. But, uh, over time, these, these things build up, and we scrutinize our false positives, our false negatives.

Elliott Margulies: Um, they’re very, very small, and they’re, they’re rare. Um, but this is a screening test that we are looking at, and, uh, those, uh, can lead to ways in which we can think about systematically continuing and to improve the assay.

Grant Belgard: Makes a quality metric useful rather than ornamental

Elliott Margulies: if it drives at, at something that is measurable and changeable. and again, that’s, that’s very much in our kind of ethos is this concept of being guided by first principles. We often want– oftentimes want to measure something directly versus indirectly. I think as bioinformaticians, it’s easy to kind of, pull together some sort of, you know, convoluted…

Elliott Margulies: I’ll call it convoluted correlation. but if you can’t detangle, you know, what are the values and metrics going into why you’re seeing this correlation, it becomes abstract and, you don’t know what you’re missing or what you’re not seeing. And so oftentimes we try and peel away the layers of the onion so that the analyses we’re looking at are rooted in first principles.

Elliott Margulies: And we’ve seen successes, in doing that. I can think of a recent example where we were doing an analysis and, you know, we could draw, you know, a weird square around, you know, certain data points and kind of increase our sensitivity, and, you know, with a minimal, you know, increase to no-call rate.

Elliott Margulies: But, at the end of the day, we actually found that, doing a more principled Z-score approach, was much cleaner. and we, we understood what was going into that Z-score, in a, in a much better way. And so, so that’s, that’s constantly on our mind as we, as we’re doing these analyses.

Grant Belgard: Where do you see AI being genuinely useful in molecular diagnostics?

Elliott Margulies: Oh, fun question that I know is, is, p-potentially talked about all the time. But even over the last six months, we are seeing bioinformaticians turn into full stack software developers, but with this scientific level of expertise where they deeply understand the workflow and the systems.

Elliott Margulies: And so they’re in some ways hyper-enabled, to move things forward. We’ve gone from, you know, building out prototypes that involve, you know, tab-delimited files and filtering them through Excel spreadsheets to now, you know, on the first try, having a customized user interface, you know, that, is designed for the particular workflow that we’re, we’re doing.

Elliott Margulies: and, that’s all empowered because we’re historically not front-end developers, but we’re now empowered to be able to be front-end developers, with the assistance of, of AI. And so I’ve seen it– I’ve seen the, the aggressive adoption, aggressive but thoughtful adoption, both as a software development tool, and also as a data analysis tool.

Elliott Margulies: So you can quickly advance a variety of different analyses, with AI assistance. We, you know, I think, again, we’re hyper-enabled because we’re already on the command line. We’re already accessing our databases and our infrastructure and our systems. So the access to the data is not a problem.

Elliott Margulies: The envisioning of what we want to do is not a problem. And so we can very well prompt and describe to AI systems what we’re looking to do. And I’ve, I’ve watched, both on the software development side and on the data analysis side, the speed with which we can move is just incredible. You almost have to pause because a human being can’t absorb it all as fast as we’re able to generate it all, right?

Elliott Margulies: And, and that’s actually, we recently had a hackathon as a team, and one of the comments was, you know, “Yeah, I built this thing so fast, but I, I still need to like pause and like look under the hood and really fully understand w- you know, what was going on here.” And, and, and, and that, that disconnect is, to the actual line-by-line coding is what everybody is starting to get used to and try and, and, and work through.

Grant Belgard: What changes when bioinformatics moves from method development into high throughput clinical operations?

Elliott Margulies: There’s, there’s a little bit more of a rigor, uh, associated with change control, um, ensuring the quality of the output is what you expect it to be. In many ways, you know, we strive for things to be as perfect as possible, but more importantly, knowing where, where it falls over and documenting that and under– and, and having an appreciation for that is, uh, in some ways more important than getting the algorithm perfect.

Elliott Margulies: We don’t– We, we never wanna be blindsided by an error or an issue that we didn’t realize our systems weren’t performing as well as they should. And, um, so we’re on– constantly on the lookout for, for that from a clinical ca- uh, standpoint, particularly also as the scale come, comes about. Um, but we have, you know, fantastic team members who think through a problem, and anytime we’re able to, um, make an improvement, one of the exciting things is that we’re sitting on, um, a treasure trove of historical data that we’re able to understand and with, with outcomes information.

Elliott Margulies: If we make an improvement, how well does it perform? How did it– How does it, uh, compare to the way our, our, you know, previous version was with, with the system? And so we, we build automation and systems that are able to essentially do an AB comparison, right, off of, you know, tens of thousands of samples, if, if not more, um, uh, to, to be able to have that rigor whenever we push a, a new change into the system.

Grant Belgard: How do you know when a pipeline’s ready to leave R&D and enter routine use?

Elliott Margulies: I mean, the, the formal way is through a, a, an actual validation. We set up a, you know, a multi-arm trial, that kind of tests both kind of individual measurements as well, you know, and then it grows from there to reproducibility to final kind of validation outcomes.

Elliott Margulies: At, at the end of the day, we’re, we’re making a, a variant call or a, an assertion of risk, and we have to have a known truth set, and we have to, show reproducibility. and then we also try and show, you know, almost like forcibly show limits of detection. Where is there flexibility in the system where even if it doesn’t, if it’s not perfect, it’s a little bit noisier data, we’re still getting the right answer.

Elliott Margulies: Like understanding that limit of detection is, is really important. But it all happens through, you know, we, we follow very standard, processes by which we verify and validate our, our, algorithms.

Grant Belgard: Now to, shift gears and talk about your career. what, early experiences most shaped the scientist you became?

Elliott Margulies: I can point back to a lunch I had as a grad student where Francis Collins was, a professor of genetics at, University of Michigan. He was now at the NIH, but he had still come back on occasion through his relationships with the University of Michigan, at the medical school, giving lectures on occasion.

Elliott Margulies: And hearing him talk about his day moving the Human Genome Project forward, is still a core memory of mine as a kind of a young, graduate student. and that planted a seed in my head about how you can take, a doctorate degree and not just be a, an academic professor writing research grants, for the rest of your career.

Elliott Margulies: Kind of the desire to collaborate on a scale that was beyond what one person can do, it started to really kind of be a seed that was, you know, sprouting in my brain. And, and then moving to the NIH for my postdoc, with Eric Green, who I know you’ve recently had on this, this podcast as well, really opened my mind to that type of at the time it was called big science, right?

Elliott Margulies: Where you could, generate data at scale, you could collaborate at scale, and it was a ton of fun. I’ve, I’ve, I’ve constantly found myself moving in, in those directions where the learning from a diverse group of scientists, is feels much more comfortable to me than working on an individual a-assay or experiment, or project o-on my own.

Elliott Margulies: I have those little pet projects here and there, but at the end of the day, it’s watching the, the, the group of people accomplish something that was not possible as individuals. You, you know, that’s, that’s been a ton of fun for, for my career.

Grant Belgard: Can you tell us about, some of the later major inflection points in your career?

Elliott Margulies: Sure. So I, I sometimes

Elliott Margulies: think about, my career, at least right now, as, having kind of three major phases, I guess. both– So I started as a faculty member, at the National Human Genome Research Institute, where I was doing more basic science research, but in this kind of largely collaborative way, trying to make an impact on the world with, this transition t-from, kind of population scale, you know, low coverage sequencing to find variation in, in the human genome to being able to sequence individual genomes.

Elliott Margulies: And how do you build the methods that can, accurately call variants in individual genomes? I then, kind of the second phase of my career is when I moved to Illumina, working in this advanced medical research group, with David Bentley, and his team, where we really thought about both kind of balancing, what is possible, and from a clinical medical perspective at scale, and, balancing that with kind of the drive of the next generation sequencing that, Illumina at the time was, was driving forward.

Elliott Margulies: and it was a fun balance, because we were trying to provide y-you know, utility to the sequence data that was being generated, and added value and thinking about, I was able in that environment to think five to ten years out, you know, how can we build collaborations to show what’s possible about bringing, you know, high throughput, whole genome sequencing to a medical system?

Elliott Margulies: You know, at the time, you know, we had recently, started the hundred thousand genomes project in, in England as, as a kind of a f-first initiative of bringing genomes into a clinical context. and then this third phase of my career, which I’m in right now, which to me, there’s a bit of, irony in the fact that for much of my career, I had spent focusing on whole genome methods for, you know, sequencing and interpreting, genomes, to now sequencing super small parts of the genome that are incredibly actionable and meaningful and doing that at, at scale, and, being so close to patient care.

Elliott Margulies: this is something that is, is incredibly meaningful to me and my team. And I don’t think it’s often you get a bioinformatics group like we have here that when we find a high-risk result, it weighs on us, and we know that we’re now supporting genetic counselors who are gonna be giving this information to families at a very, you know, that are, the, the, at a very challenging time in their lives.

Elliott Margulies: But they’re going to be better for it because they have this information. We really believe that having that information is, is empowered, empowering to families. but we are, as a bioinformatics group, so close to those decisions, and supporting teams who are delivering this information. It really is remarkable and, in other parts of my career, even though I was building tools that were interpreting genomes or building methods that could identify variants in rare diseases, I was never this close to patient care, and it’s very rewarding, to be able to, to support an organization, like this.

Grant Belgard: How has your view of sequencing changed as the field has matured?

Elliott Margulies: Well, I think we’ve reached a point where, uh, sequencing is a, a commodity. Um, I mean, I r- I remember a time when it was very novel to be able to do massive short-read sequencing. Um, that is now a commodity, and it is just one part of a toolkit that enables clinical adoption. Um, and something that I’ve learned a lot from this most recent part of my career is how much do you have to think about holistically, a, a sound business model that is reimbursable and you’re building a system that, uh, can operate at scale, uh, within certain cost measures and where sequencing is just one part of, uh, of that cost.

Elliott Margulies: Um, but thinking about how to interact with insurance companies, with physicians, um, with patients is all part of the process now. So it’s, it’s become… It used to be a, I guess, to, to answer your question, it used to be a major part of everything that I thought of, uh, when– as a bioinformatician, and now it’s just one part of, of many that are equally, if not more important, because we’ve essentially succeeded in making DNA sequencing a commodity.

Grant Belgard: What have different work environments taught you about how science gets done?

Elliott Margulies: It’s all about the people and the relationships. I’ve been around smart people I think all throughout my career, whether I was in grad school, as a postdoc, um, at Illumina, here at Billion to One, and the times where we were successful and impactful were the times when we could get large groups of people working well together and feeling comfortable and open to uh, making decisions under difficult situations, not being afraid to call out when there are issues or concerns or timeline risks or technical risks.

Elliott Margulies: Um and I can point to the exact opposite happening, where, um, sometimes great ideas didn’t succeed because people were afraid to say, “I don’t think this is gonna work,” or, um, “I don’t know if this other person agrees with this, this plan forward,” right? And you know, everybody can be well-intended, but, um, if you don’t have that open, uh, way of communicating, uh, and being able to, um, talk about the, the more challenging things in a calm, open way, um, uh, oftentimes you can’t see these, these great wins and these great successes.

Elliott Margulies: Uh, and I found a company that truly values all of those, those parts of the, the scientific process and the development process. So I really feel supported, both w- my kind of senior colleagues, as well as the team members that, uh, I’m mentoring, um, to be– to foster this type of environment.

Grant Belgard: What’s something you’ve learned the hard way?

Elliott Margulies: I think finding an organization and a mission that is aligned with your own career goals and growth is important. I can think of a time in my career where I thought I was well-aligned with where I was trying to move things forward, but it ended up not being the same as the organization that I ended up being in.

Elliott Margulies: And for a period of time, I really struggled. Not because somebody was, was trying to set me up to fail, but because my interests, uh, and where I thought the organization wanted to go was different than where the people who were leading the organization wanted the organization to go. And, uh, it took me a while to realize that disconnect.

Elliott Margulies: And once I kind of realized that and I was fortunate enough to be able to kind of move into a different organization that was well-aligned with my goals um, my career started moving forward again. And I think that’s, um, often-oftentimes, right, nobody’s ill intent, but, uh, when you get that mismatch of, like, where you want to go, what your passions are, and if that’s not aligned with the organization that you’re in, um, that can be a rough, rough time.

Elliott Margulies: But to your point, I can also look back on that and say I learned a lot during that time that, uh, I wasn’t necessarily move-moving my career forward. Um, in hindsight, it has helped me move my career forward because there was a lot I learned during those periods of time.

Grant Belgard: Related to that, have mentors taught you that still shows up in your leadership style?

Elliott Margulies: I can point to very specific things that I do that are relics, so to speak, of things that I learned because mentors have done them to me. And some of that is fostering that deep collaboration. So Eric– I can think about Eric Green, uh, working in his lab. He just opened doors to, for me to collaborate with people that I didn’t, uh, know I could collaborate with, and I learned so much about that.

Elliott Margulies: And I enjoy when I can bring people of kind of from different parts of the company together and start to see them, their minds just go like, “Oh, wait, I’m thinking about this problem too.” And, and they’re thinking about it from a different angle, and you kind of see that, that collaboration take place. I can think about, you know, as I move my team forward, um, and I constantly am trying to think about how are they…

Elliott Margulies: Are, are they happy? Are they growing? Are they doing what they want to do? Uh, w-when we get a chance to be together, recently we had our, our teams off-site, um, we all shared a success and a challenge. And giving everybody the space and the comfort to share a challenging moment that they might be going through with the, with the entire team in a way that is very healthy and productive, and they don’t feel like they’re going to be shamed or, um, or, or disadvantaged because they shared something they’re struggling with.

Elliott Margulies: To create that environment, it becomes very empowering, and that’s something that David Bentley taught me. Uh, you know, so I can think about, you know, that, that type of interaction of leading a team that’s very thoughtful, um, and constantly trying to grow and these are things that, that, uh, definitely come from great mentors that I’ve had in the past.

Elliott Margulies: On a lighter anecdotal note, uh, music is really important to me. Francis Collins is somebody who always brings music to a lot of the science that he does and the way in which we kind of bring teams and communities together. Um, I started my off-site by singing a song to everybody and having them sing with me.

Elliott Margulies: Um, these are things that I, I often do that are fun to think back on how wonderful leaders and mentors have kind of touched my own life and, and career and how I’m trying to give it forward to the, the rest of the team that, that I now lead.

Grant Belgard: What do you look for when hiring bioinformaticians who will thrive in clinical diagnostics?

Elliott Margulies: So we hire people at different stages of their career. I’ve been very impressed with the tenacity and thoughtfulness of the right person who is kind of just coming out of an undergrad or a master’s degree. The work that we do can, be very impactful and has to be done in– with, with a lot of thought, because these are patients who may be getting a high-risk result.

Elliott Margulies: And it takes a certain kind of individual who has a level of maturity, to, understand that what they’re building has to be right as best as they can do it, where you have to own up to a mistake, in a, in a way to– that can be productive, where both you learn and we learn how to improve the system.

Elliott Margulies: so I think finding a smart individual who can code and run bioinformatics pipelines is kind of table stakes. It’s really looking at an individual, what are they passionate about? How do they wanna make an impact in the world? How do they wanna grow their career? Are those passions aligned with being in a company where, we are managing patient samples?

Elliott Margulies: being in a company where we’re doing science that is leading to new products, where we’re going to be balancing an intense period of time of scaling and b-and building new products and trying to automate more, so that multitasking ability. But at the same time, tr-tr-uh, trying to ensure that people can focus and not get overwhelmed, with everything that’s going on.

Elliott Margulies: So these are, these are kind of traits we try and, and look for. you know, these are insights that, that go beyond just coding. But, you know, can you, can you analyze a data set and, you know, describe what you’re doing while you’re coding and, you know, present the, the results to people, in a meaningful way that are, you know, getting back to one of your earlier questions, you know, that the cross-collaboration, is key.

Grant Belgard: What advice would you give to computational biologists who want their work to have clinical impact?

Elliott Margulies: Well, definitely align yourselves with, a company or a lab that has a clinical focus. I think more and more, we’re seeing the ability to translate genomics into actual clinical results. And, it’s not, you know, to something we were talking about earlier, it’s not just a matter of sequencing somebody’s DNA and doing some variant calling anymore, right?

Elliott Margulies: I mean, there’s a whole system and an organization of how do you interact with patients? How do you interact with healthcare systems? How do you interact with reimbursement? and being able to… Something that I don’t think I fully appreciated till I joined Billion to One was that this organization has that all, and this is why I think I’m feeling empowered to succeed bec- in a clinical context.

Elliott Margulies: Whereas I thought I was doing clinically impactful work, and I probably was, right? But it was more futuristic research. But what I’m doing right now is, is impacting patients today, and, that’s very, very meaningful. So looking for an organization that is thinking holistically about all the different aspects of molecular diagnostics, not just the molecular part of molecular diagnostics.

Grant Belgard: What should trainees spend more time learning than they currently do?

Elliott Margulies: I think understanding a lot of the science and engineering principles is b- is going to be incredibly important because coding as we know it today is starting to go by the wayside. it is somewhat of an identity crisis for bioinformaticians. but it means that understanding how software systems are architected, that’s not gonna go away, right?

Elliott Margulies: you still need to guide an AI to build the right tool in the right way. and making those decisions, requires knowledge of how to build scalable software systems. There’s a kind of an architectural engineering component that I think is gonna become more and more important, as well as the science, of things, right?

Elliott Margulies: B- being, you know, being able to, continuously learn and adapt to the science of what you’re doing, will be essential, in, in moving things forward. We’re, we’re no longer just people who plumb and build pipelines. that’s not what we do anymore.

Grant Belgard: If you could go back and give advice to your younger self, what would, would it be?

Elliott Margulies: don’t underestimate serendipity

Grant Belgard: And, to, to wrap things up, are you most optimistic about in genomic medicine over the next decade?

Elliott Margulies: Finding deleterious pathogenic variants and, medicines that can overcome those deleterious mutations, that is an incredible change where we, you know, first we had to develop the tech to identify mutations. Now we can identify mutations. And soon, not only can we identify these mutations, but we can proactively provide pharmacogenomic, you know, information, and, and drugs that can mitigate, those mutations.

Elliott Margulies: That’s just an incredible place to be. We’re seeing this in the prenatal space now, start– just starting to emerge with a couple of key examples. but that’s, that’s a wonderful place to be because now you can kind of complete the cycle of not just delivering challenging news, but, ways in which you can mitigate, those, those mutations.

Grant Belgard: Have any, any closing advice for, career scientists and bioinformaticians listening to the show?

Elliott Margulies: Focus on what makes you happy, uh, you know, and, and curious. Uh, stay curious. I think, um, it’s important to realize that what you learn today is not necessarily what you’ll do in five or 10 years. But maintaining that ability to constantly learn and adapt, uh, is, is a, is a fun place to be. And again, you know, lean into the serendipity part.

Elliott Margulies: You, you never know, uh, when, uh, you’re gonna work on something and the– this particular project is gonna be the extra meaningful thing. Uh, you didn’t plan for it, but it’s turns out that the next three months are, uh, uh, an, gonna be an intense moment where you might learn an extraordinary amount in that short period of time that you didn’t plan for, and you’ll make a, a disproportionate impact by investing in, in those, those periods of time.

Grant Belgard: Elliot, thank you for joining us.

Elliott Margulies: Oh, it’s been my pleasure. Thanks for having me, Grant

The Bioinformatics CRO Podcast

Episode 86 with Deniz Kavi

Deniz Kavi, CEO of Tamarind Bio, discusses building a company and how to make great software for life scientists.

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.

You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Deniz Kavi is co-founder and CEO of Tamarind Bio, which provides accessible molecular design software for life scientists.

Transcript of Episode 86: Deniz Kavi

Disclaimer: Transcripts are automated and may contain errors.

Grant Belgard: Welcome to the Bioinformatics CRO Podcast. Today, I’m joined by Deniz Kavi, co-founder and CEO of Tamarind Bio. Tamarind is building software that helps scientists access computational biology tools at scale through a web platform and API across areas like protein design, structure prediction, and docking.

Grant Belgard: We’ll talk about what Deniz is building now, how he got there, and what advice he has for people trying to build at the intersection of biology software and AI. Welcome.

Deniz Kavi: Thanks for having me

Grant Belgard: Thanks for coming on. So tell us about Tamarind

Deniz Kavi: Yeah. Tamarind as you described, basically a aggregator of all the leading molecular design tools. So things like AlphaFold and RF diffusion and many others, we provide in a single place for all scientists to use, basically. So we try to do the work, like the boring, annoying, undifferentiated work is what we try to do our best in.

Deniz Kavi: So yeah, scaling up these tools, applying them to real problems, connecting them into multiple stage pipelines, so on. I got into this problem basically doing this as a human for my colleagues at Stanford. So in my undergrad lab, my job was to get an email, run AlphaFold 10,000 times, email it back to one of my colleagues, and at some point I felt there does not need to be a person doing this.

Deniz Kavi: You can just have a software solution from there. So basically trying to lower the accessibility gap, but also handling the infrastructure, the, software aspect of how you deploy computational tools like molecular design models

Grant Belgard: What are you most excited to build or learn about right now?

Deniz Kavi: Yeah, I think today we spending a lot of time on developer tooling to some extent. I think we have found that sort of– we work mostly with large pharma and biotech companies, and we found that these organizations have a lot of interest in how they do– how their R&D teams deploy AI to their use cases. And, ultimately these have served that demographic and like we are building a lot of tooling for that. So I think most exciting– most of what’s most exciting to me right now is a organization-wide, whatever, one to ten thousand person deployment of molecular AI tools for, serious drug discovery applications, things that will go to clinical trials or become tool proteins, what have you. Yeah, I think serving data scientists and AI folks to then deploy that to the whole organization is something I’m very excited about today.

Grant Belgard: What problem do you most want to solve for scientists?

Deniz Kavi: I think the general genesis of the company has this theme of accessibility. Historically, if you look at how software has worked in our industry, do a– you do a computational chemistry PhD, you’d take a four-month course on this one software product, and you become an expert on it, and then you are the person tasked with running all these workflows for your colleagues. I think we think there’s a spectacular amount of inefficiency in sort of the handoff in information or in convincing your computational colleague to do this work for you, like all the lobbying internal work. I think our general case is that of the increased capabilities of AI recently, it’ll be available for every scientist and not just specialists in, 20, 30-person data science teams, but the entire org will be the consumer answer of people applying the AI tools in our day-to-day.

Deniz Kavi: So I think very much the, having every scientist be empowered by computational tools as opposed to a realm of specialists doing it for them

Grant Belgard: What’s the backstory on the name?

Deniz Kavi: Yeah, unfortunately not a very exciting story, but we basically wanted it to be named after a living thing. We went through a list of trees, found tamarind as a, memorable fruit and a notable name, and then went from there

Grant Belgard: Where do you think the real bottleneck sits today? Is it model capabilities, infrastructure, workflow design, adoption, something else?

Deniz Kavi: Yeah, I guess I would be biased to say infrastructure. I think model capability is a credible thing to look at as well. I think the better the individual tools are, the more– the less certain-uncertainty you have, the more reliability you have on the individual models, the sort of more valuable the infrastructure becomes.

Deniz Kavi: Because if you have to tweak a bunch of knobs and find all the errors and evaluate against them, the model itself may not be very useful if you don’t know those exact details. But in our case, we’re seeing the molecular design tools especially become good enough in the space that, you can just point at a problem and apply it there and go from there, basically. Our general view would be having it be available to a scientist. I’m repeating myself over and over. So generally, very much the sort of focus of the company is how do you get somebody with an interesting idea for a target or a, drug program or a way to optimize a lead protein or lead sequence and how do you get them to convert that idea into, a molecule on the computer, basically.

Deniz Kavi: I think broadly the issue is not so much the hypothesis generation or having targets to go after. It’s more so like how do you– Once you have a problem to solve, how do you apply tooling to solve that problem? It’s like how we think about the, one of the gnarly-est problems in the space.

Grant Belgard: What does great software for life scientists look like for you?

Deniz Kavi: Yeah, I think we, from the start, focused on having an opinionated way to apply these tools in the first place. You– if you take a look at the papers of these there are a lot of focus on, like, how the model is trained, how their– how the inputs should look. Our view has always been basically, we built a sort of opinionated, structured way to consume those tools in a serious, drug discovery use case. So for structure prediction, complexes and docking is a big use case. For protein design a lot of papers come out and say just, we are creating a way to create new structures, and these are arbitrary, not really– these are not becoming drugs, these are not becoming antibodies or vaccines or what have you. And we build templates and recipes to make that work more effectively. That’s one avenue, is like you need to codify and build the best practices into your product.

Deniz Kavi: The other side is, I think, integration. So we have been very like the entire product from day one has been usable programmatically.

Deniz Kavi: So our API or MCP or agent offering is basically a way for any arbitrary code to call our product. So I think in a way, we can just go away in the background, just be done as infrastructure, and that just gives you flexibility to consume the tooling in wherever you want it. So it might be in your ELN, your LIMS system in your internal custom applications or your AI agents within Claude, what have you.

Deniz Kavi: I think those two are the main points of you want to build a opinionated, codified way, like a very practical-focused applications of the AI tools. And the other side is the how do you get it to– how do you get a scientist to use you where they want you to use them? That’s the two pillars, I would say.

Grant Belgard: What part of the scientist experience do people outside the field most often miss?

Deniz Kavi: I think amount of, uncertainty in everything you do, basically. Like I think some people think about the process as, I have an idea, I will do an experiment to fail or not fail that idea, and then that’ll be the end of it. And if it works, it’s gonna be a drug and it’s going into clinical trials.

Deniz Kavi: If it doesn’t work, I like toss it. But oftentimes, there’s uncertainty in the actual experiments you do and the computational tools you use. ‘Cause anything you do has some error bars associated with, and you have to believe and persist every time something is like going wrong where it might be a sign of it being interesting. Very rarely something in science is, immediately a useful result or like you, you even get data that like proves it in either direction. So I think that’s a very much something people not think about is you’re not really sure where you’re going when you’re in the middle of a process or like hypothesis in general.

Grant Belgard: What have conversations with users taught you that theory never would have?

Deniz Kavi: I think we don’t talk much about– People talk a lot about like where AI is interesting, where the applications are, but applications like are pretty broadly on what people think about as things that computers can do. I think if you think about the grand idea of a virtual cell, the sort of like a philosophical leaning of a research work there would be you want to create a thing that simulates the entire behavior of a cell in the computer. But realistically, that is like too big of a problem to be applied immediately today. So in some ways, I don’t have– I have a computer science background. I have a sort of, basic science background in, in relevance to that as well, but I’d never worked in a sort of specific R&D org before in a drug discovery context.

Deniz Kavi: And I have learned a lot about what the day-to-day problems are in a specific, in the context of a research program for a drug program what the use cases there are. So that’s things around what assays are most interesting, what sort of you can trust on, what you can trust what the sort of clinical or risks are before going into the later stages.

Deniz Kavi: I think that type of day-to-day experience is very hard to see. Certainly in academia, even if you’re a very strong scientist, you might not know the sort of how the applications go. And there is a meaningful difference in academic science and, commercial drug discovery-focused science, I would say. And, having not worked in that space, talking to– the privilege of talking to like pretty senior folks who manage these R&D programs and learning how they think about those use cases is very valuable for me.

Grant Belgard: How do you decide what to hide behind the interface and what scientists should still control directly?

Deniz Kavi: That’s always a changing structure for us for the most part. I would say people tell us that this point is a big value add for the platform, like maybe whatever tens of thousands of scientists, people will email us when they want a model added. I think we’ve always had a question of complexity versus abstraction, where, we have, say a few hundred person org in a company using us and also like some 30, 50 person data science team using us. And the interests of those groups are not necessarily the same. The scientist wants it to be easy to use. They can come in, use it, and then forget about it, but the data scientist might be, I want to check every single possible option. Generally, we’ve built our product in two distinct sort of ways.

Deniz Kavi: So the API, for example, is you can– you do anything you want programmatically, can write code to build on top of our product. And then nowadays, the API actually serves as a way to like abstract even further. So we found AI agents as a good way to guide the user to, you know… Most of our value from AI agents, like not even necessarily the reasoning, but just the hiding the complexity of how you consume a tool.

Deniz Kavi: That’s the one avenue. And the web interface is also there for the most complex way to Basically, anything you need for a given tool, you can use with another web interface. So we bifurcated our products in that way for different cases. And then generally just being close to the customer, and people tell us when they’re unhappy and, it’s a boring answer, but just people– you be very close to, your check-ins and regular meetings, what have you, where people will tell you what the, what problems they have and how– whether or not we’re able to solve them.

Deniz Kavi: Or if we can solve them, are they able to access that problem or access that solution to that problem are the main ways.

Grant Belgard: What does a trustworthy AI-assisted workflow look like to you?

Deniz Kavi: Yeah, I think generally the– I guess these are two different answers for, what we live in, the molecular domain, and also the sort of general knowledge work reasoning LLM tasks. Would say the molecular side, you just want wet lab data. I think we’ve seen a big explosion of the recent models doing very well in the sort of classical, one binder design or structure prediction problems, and much of that was launched by you can do experiment to see that this model works and like you can see the hit rates, what have you. I think that both added a lot of trust, but also despite not substantial changes in architecture, showed just that the models can be good and you can demonstrate that they’re good in that way. I think there’s still some questions there, like how much you can see from experiments. There’s– obviously, you can’t run a clinical trial to show that your model works every time.

Deniz Kavi: The molecular design side, I would say it’s the, just want to run experiment that can be as close to the use case as possible. Then surprisingly enough, like in some, most of the best performing models, they don’t really tend to correlate with a direct experiment. Like most of them are not simulations of an assay you would do.

Deniz Kavi: Many of them are, but not probably the most successful aren’t. That’s one. For LLMs, I think the

Deniz Kavi: main

Deniz Kavi: question is around, You know, to some extent, the tools LLMs have access to are not particularly strong. I think if you give the tools a bioinformatician uses to an LLM, there’s still some uncertainty and problems around that because of the sort of inaccuracy of the specific tools you’re using or the value of the, the reliability and the accuracy of the tools it’s using in the first place. If you give an LLM a protein design model, you now have two things to worry about, which is a faulty tool using another faulty tool, and you have to make sure that you have to find out which one’s the wrong part there. View for there– for that has been basically being very transparent and clear about what the tools it’s using are looking like and evaluating those results.

Deniz Kavi: Our LLM agents basically call our same– Like, whatever humans would use, the LLM uses that tool, and you can interpret it in the same way. So all the sort of inter– the abstractions we built help there. other questions are how good is the LLM in scientific parsing and understanding what’s going on. at this moment in time, the, Claude Op– Claude 4.6 and as of recording is 5.4, Thinking for GPT are reasonably good at not hallucinating anymore. I think questions tend to be where do you take the result of a paper at face– like what the result you’re citing at face value versus is it applicable?

Deniz Kavi: So let’s say if a model setting use a tool, the only tool it has access to for, let’s say protein solubility prediction, it just says protein solubility predictor, and it does not have any error bars or caveats or any details there. The model has no choice but to just, “I have this. This is the one thing I have.

Deniz Kavi: Let me just call this and give it to my user.” But in fact, maybe refusing to do something or showing its limits is probably a most more valuable thing there. Somewhat long-winded answer, but I should say, and showing uncertainty and then maybe even rejecting tasks when it’s on not confident in its ability to do them are two major points, I would say. And so I think the RLHF efforts on the LLMs are doing this well for other– medical advice or what have you. I think these will apply for science because we’re not really– we’re not trying to adversarially convince the model to do something that’s bad because there’s no benefit to having it do a ineffective in silico experiment, I should say.

Grant Belgard: How do you think about simplicity without flattening the science?

Deniz Kavi: Yeah, definitely a difficult problem. I would say the honest answer is like education and support is the honest or best way to go about this. I think we, we try not to limit capabilities of the tools we serve. To some extent we do, honestly, like for the, again, the, say small molecule design tools, for example. Like so many custom options you could be doing that it becomes untenable to use them for anything in the first place. So we try to build recipes of best, best– most commonly used use cases. I think the, again, the sort of open-ended AI models, like things like, protein design or small molecule discovery or even the LLMs themselves, they have a lot of use cases that can be disc– discretized into buckets of, this is my one use case.

Deniz Kavi: I want to do, docking prediction for small molecules, and I have a big world model that does it for me. I think narrowing those down to different tasks is the current approach people take. I think this is like unfortunate because if you can have a single model do everything, it’s quite, quite a more interesting experience and more interesting results most likely. But for the most part, it’s carving out chunks of the capabilities and giving them as recipes you can use is one, one approach. So you don’t really use– lose the capability there. You just become a bit more constrained into a way of doing things, which can be good if the model is not that good, is the– not that performant is the other detail there.

Deniz Kavi: For the most part, I think focusing on the where the models are good and like you can know that by being close to your customers, but also being close to the results in the first place with a sure review and then carving out those places to be individual tasks you’re looking at are the things we try to do.

Grant Belgard: What things are easiest to overbuild in this space?

Deniz Kavi: I think you can make a lot of analogies to like horizontal software players. So I think in some ways we resemble something like a Databricks or Snowflake quite a bit. We do model inference, training onboarding our custom tools, app builders most of those hold, but those products are also used a lot in the life science industry, and they can’t quite apply to them directly. the over-index on verticalized solutions for the industry. You think of yourself as a very bespoke product. There’s lawyers who do life sciences, there’s, procurement teams who do life sciences. And in fact, some of that might just be if you just take the horizontal version, like up and modify and optimize, that might be a decent way to approach it in the first place. You don’t need to reinvent everything.

Deniz Kavi: You can probably use the, the workflow builder, the pipeline builder for generalized use cases probably works well for the bioinformatics use cases as well. I think we don’t need to overindulge in our specialness and difference from the software industry or whatever the AI applications are.

Deniz Kavi: I think that’s the one thing I see a lot of is, you want to reinvent the tool hosting and say the bio– bioinformatics is basically, computer science, and then you don’t– a lot of the tools there, you can apply to the general use cases. So I think we– both from a people who gets into the industry perspective, like we hire a lot of non-biotech folks on the software side and then also what tools you use to build products are the– they don’t all need to be from scratch.

Deniz Kavi: They can be building on the horizontal solutions.

Grant Belgard: When new tools or models appear, what makes you pay attention?

Deniz Kavi: Yeah, I think the– surprisingly enough, I think the, to some extent, like the social media hive mind is like a decent predictor of like how much interest there is in something. That doesn’t necessarily mean it’s like the most accurate or the most interesting. Like some of our models are probably very, like very well known.

Deniz Kavi: People outside the industry know them, what have you, but they’re like basically not used. I think some level of interest tends to be a sign of in- of value. I think if something just goes out to bioRxiv and nobody talks about it, like it doesn’t tend to be the most valuable. I think that like in that case, like the most obvious thing is immediately valuable. The other side is these days, a lot of validation you can discretize and show for, especially in our use cases, in random molecular design use cases. I think if you show clear objective results there I think we’re doing a decent job of putting out benchmarks. There’s more to be done there for sure as well, like for other use cases. In my mind, it’s basically two, two avenues. One is, it can do something that was never doable before. I can make a fully– a protein completely from scratch and put it out and express in the real world.

Deniz Kavi: That’s meaningful scientific results. The other is it does something that was s- theoretically doable, but it was so much better than that, say 100 times better than that, that it can reliably get there.

Deniz Kavi: So again, going back to the protein design example, we’ve had a renaissance of binder design tools recently, and those are mostly because they’ve had, 100 times the hit rate of the previous generation of tools, and now people can actually use them and, instead of, they don’t need to have a huge V display set up where they spend a million bucks on the assay.

Deniz Kavi: They can just spend, whatever, $10,000 and show the results there. So I think it’s either increased accuracy and then the other side is just a net new use cases like in-invented by this protocol.

Grant Belgard: What trade-offs do you face between frontier capability and everyday reliability?

Deniz Kavi: I think you don’t want to be overambitious in applying tools to a given use case. Theoretically, there’s a model that claims to do everything you need to be doing for a drug discovery program. If you read the Wikipedia or you said the GitHub repo the definitions, and you just did them, you say you gave them to an LLM agent and then told them, “Don’t question these.

Deniz Kavi: These can do exactly what they claim to be doing.” And then probably like it will fail in the first step, but it’ll fail in the second step and the third step and fourth, con– successively. I think there’s a– incentive is always to publish like a decent result, and you just claim a… Maybe the most humble people say towards this task or towards, towards even a Jesse prediction on the computer. You do need to build the muscle of identifying what is there to publish a paper or what is there to show an actually or what shows an interesting result. I think that’s one question we have. Something that might be at the frontier might not be good enough for it to be actually practically useful.

Deniz Kavi: GPT-2 was on the frontier for a while, and it was basically useless for a lot of tasks. And similar things apply for our world, which is basically the incentive is to provide the most audacious result you can. But even– That might be true in relative terms. It might just be, it might be 100 times better than the previous stage, but that might not be enough from– If it never worked and it’s zero is zero.

Deniz Kavi: So just making sure that the individual applications are– I think restraints are, I should say in that note, restraints are a big thing we think a lot about where can a model be applied, what the actual sort of error bars around the results are, how can we codify that into a product where you can understand what the limitations might be. I think it’s just that there’s a lot of things that claim to be the best, and then the best is not good enough for a lot of use cases too.

Grant Belgard: What do outsiders misunderstand about making advanced computational biology tools genuinely usable?

Deniz Kavi: I think the view or the understanding is basically that everybody’s aware of what the protocols they want to use are and like what the best, what the hit rates are, what the– People think about, computational modeling tools as the same as they would think about like assays, where there’s some clearly defined set of worlds.

Deniz Kavi: You can buy a recipe or have a recipe built internally and just apply that, and like it’s, 90% of the time it’s gonna be the same thing, except some modifications you have to make, what have you. I think the– both the world is changing so quickly that like very few scientists are actually on top of what’s the frontier.

Deniz Kavi: I think like we do that role informally by just giving guidance on that and providing educational materials there. For the most part, yeah, the limitations are just, people don’t really know what works as a community, basically. So like I think it’s very much in flux, and that might maybe be stabilized like every couple of weeks and then we go off the rails again. The question basically is: scientists know what they want to do? Do scientists know what works well? And, the criticism of the science community is a sort of general observation that the world is moving too quickly and the incentives are, or incentives are such that everybody claims state-of-the-art for everything, and it’s hard to be on top of what the direct applications are.

Grant Belgard: What first pulled you toward the overlap of software and biology?

Deniz Kavi: Yeah, I actually– so coming into Stanford, I had done some work for doing natural language processing work. I very wrongly predicted that sort of transformers would not be a good application for life sciences. And so I was doing some older text translation and summarization, sentiment analysis type work, and I just assumed, “Oh, these are not gonna be useful tasks,” and, I’m gonna look at another application. And I got into biology from trying to find cool applications of AI tools as I was going into a pre-science degree. And in reaction to that, I got a job in a drug discovery context at a lab at Stanford. Yeah, I think it was very much I was trying to find an interesting application for computer science, which sort of led me to biology.

Deniz Kavi: I’ve heard this from a lot of mathematicians as well, where the modeling of a complex system tends to be an interesting problem as a mathematician. You become a bio-focused mathematician or what have you.

Grant Belgard: Was there an early moment when you saw a gap between what scientists needed and what the tooling actually allowed?

Deniz Kavi: Yeah, I think the most obvious for me personally, the earliest sign was basically that I was hired for a full-time or full-time job doing this for my lab mates. Somebody– I think to this day we see a lot of this as well, where even large pharma– A director of data science at a big pharma company who has 15 years of experience, half their job is just somebody sends them an email and says, “Can you run this in silico workflow for me?”

Deniz Kavi: And then they do it for them, but it like takes a few weeks, what have you. I was basically doing that for my job. So we had a protocol for finding targets for peptide candidates in a, in silico way, and that was interesting for applications of my lab, and that sort of led to some other separate work that was a cell systems publication. And that was out there in the world for people to use, and then nobody in the world could use it, basically. They would have to have me do it for them. And that was the genesis of the company in general, was just solving that problem of, I have an idea to apply this, but I have no idea how to actually, use the tool to apply them to a real use case. And that’s how the company came to be too.

Grant Belgard: What did being close to research environments teach you about how science really gets done?

Deniz Kavi: I think that there’s just, the most part, being close to scientists makes you appreciate both how valuable their work is, like how much how hard they’re working for these use cases. But the other side is, I guess I’ll– going back to my previous point, like a lot of uncertainty around how well do you think your idea will work. The experiment you do might fail, the computational tool you use might fail. There, there might be some, side effect you’ve never known about and, that sort of– you have to embrace that as a researcher. And I think researchers also think about the world in this way. I think it’s more, more importantly than typical day-to-day experience is a very strong skepticism of everything that comes their way.

Deniz Kavi: Like anything that could be interesting or useful, people are, just by training of the job, I think very much skeptical and often negative on these things, which I think is an interesting habit to have, and I think it has its pros and cons. I think the temperamental approach is like what the temperament of a scientist is, or life sciences, I should say, is for what the future holds, quite interesting, where they simultaneously have a job that is much, I’m in the frontier.

Deniz Kavi: I’m trying to invent the next thing, but also I am also very skeptical and generally negative on everything that might be interesting. And you have to fuse these ways to do the process of science basically which is– I thought was quite interesting.

Grant Belgard: What made the founder path feel right for you?

Deniz Kavi: Yeah, I think in our case, we built the first version of Tamarind for my lab as like a side project with my co-founder basically, and gave it to my colleagues. And then I posted on some random forum about, we have this use case, these alpha fold on the web, and web, and that led to several hundred users.

Deniz Kavi: So there was a lot of demand for academic users to– who like felt this in– problem of how they do computational tooling. that sort of felt strong enough as a pool of interest from users that we felt we, we should do this full-time and not spend our times doing this as a side project in school. And I think that was the primary reason. Like for the most part, in my mind, always not really thought about being– Like I always wanted to be a scientist or a researcher, and the approach in my mind was: What is most valuable to the world right now of the things I could be doing? What is the fastest path I could get there? And I also was very much pulled towards this idea of, I can move faster in a quick, quicker way than typical timelines you would take.

Deniz Kavi: Instead of doing a– four-year undergrad and a PhD and another, postdoc and many years of working at a biotech company, I now get to basically do the job of a executive at a relatively serious, bio-software company and just just by having the virtue of having built this thing as opposed to getting approval from everybody else on top of me is the sort of primary motivation.

Deniz Kavi: So combination. We had a thing that was working that we wanted to make it more serious. Just the waiting and the having to play the games of getting, accredited and approved and going better and better over time by some other institution was not very appealing to me. And the other side was basically for us to realize this idea of how we get this to– these tools to be used by scientists more effectively and at more practical use cases. It do- it has to be a company as opposed to a nonprofit or academic research project because the incentives are not strong enough for those use cases or for those methods of applying this problem that we had to make it a company, basically.

Grant Belgard: How did you learn to tell the difference between an interesting technical problem and an important user problem?

Deniz Kavi: Yeah, I think I was somewhat commercially minded which I didn’t realize at the time, but I’m basically a salesperson now, so I’m more aware of it today. I think many interesting technical problems can be commercial or user problems as well. The question is, we’ve been asking for people to pay us from day one, basically, as the one side. I think if you keep it as a free tool that everybody can use, it’s a– have too many nice-to-have type problems. I think you want a investment from the user of their time or their money or their feedback, what have you, where the problem has to be severe enough in their day that they would take some risk, basically, is how we think about that. I think I have a decent way to emulate that in my head today. I can understand what the workflow looks like from just talking to somebody about their day-to-day.

Deniz Kavi: But for the most part, do inference scaling for AI models. The condition for that to be useful for users is that the AI models are useful for them, and we had to validate that for that use case. my finding was, if you are willing to put something of yourself on the table, like risk something of yourself or time or money, what have you, is the main beneficiary of user solutions. And then technical products are often, you know– I think engineers are very much drawn to this case of “I think this is interesting. It’s a thing to solve,” but there’s no real thought about what it means to solve a problem as opposed to research curiosity. Yeah, that’s my general approach is getting commitment from the user is the main thing to look at.

Grant Belgard: What surprised you most about turning research-driven insight into a company?

Deniz Kavi: I think I was broadly surprised by how disconnected academia is to some extent from industry. Like theoretically, all the IP from these companies, all the biotechs comes from industry. It gets spun out of labs, it’s the professor’s research, what have you. But in many ways also, there are dedicated industry research conferences.

Deniz Kavi: The problems you tackle are quite different. Yeah, I think a lot of the basic science is sort of– you know, obviously, there’s some actual practical R&D work also happening in, in academic settings. I think very much the context was, biotech and pharma think about the world in a, series of drug programs or like potential assets you can acquire versus, research will be more platform style.

Deniz Kavi: You want to get a– I want to make the machine that makes drugs as opposed to I just want to make a drug and get– create a company around that or, provide a way to create that drug in the first place. Generally, I would say industry is much more open to investing resources to make things more convenient and efficient versus, academia or research in general can be, I will have an undergrad do this for five, 10 years and just keep going through and churning through them.

Deniz Kavi: I think there’s a lot of benefits to some problems that can just be like thoughtfully thought of and then ground through for many years or many long amounts of time. But it also means the timeline for everything you do in academia is very much shortened by what you can do. Like academia or research in general can be a source of commercialization. But if you approach the world from a perspective like I am doing something that’s very hard, it’s impossible to do, and I’m just like seeing how close to impossible I can get, that sort of somewhat grounds you in your ability to do interesting sort of groundbreaking work in some way versus if you think the thing you’re doing is very hard, but it’s like possible and like you have a timeline for that.

Deniz Kavi: I think, projects that are set up in a way where it’s gonna be I’m gonna do this in the next two years versus I’m gonna do this however long it takes tends to be an interesting dichotomy there.

Grant Belgard: Where did you underestimate the human side of the job?

Deniz Kavi: I think for the most part for us it was adoption. So there’s, one– even once you like c-close a contract or a customer signs up, like you use it within their organization there needs to be a deployment process. You need to be able– can you do basically lobby and convince functional heads, get their teams to use us? More than I… And think we have a more rational sales process than we most other industries do. Like we don’t do these steak dinners or whatever. Like it’s mostly a case of proving product value and then scientific interest and value there, and getting people to trust you and be the source of expertise, I think pays a lot of dividends.

Deniz Kavi: And that is to some extent bounded by knowledge, but also the cases, you want to be writing a lot, talking to people a lot, and that sort of tends to be the way you become a brand or a known entity in both attracting talent or customers or what have you, is very much driven by these human social problems. And then I would say in the deployment side, it’s very much driven by like political problems, like how you get a team to adopt you and how do you get a larger team to adopt you and how do you get them to use more or pay more and what have you are… ideally, you’re convincing that they’re– you’re spending more on Tamarind is a sign of getting better science, but also more adoption means more people have access to these tools that are more practically useful for their day-to-day. I think I’d never really done any of these tasks.

Deniz Kavi: I think I’ve become a salesperson, as I mentioned, and learning that on the job was very exciting.

Grant Belgard: Looking back, what moments feel like the real turning points?

Deniz Kavi: I think we– So the company’s about two and a half years old today, or two years and a few months. And the first year, we basically didn’t really know what we wanted to build. We had this idea of a, web interface for all the models, and you can use them easily. But the models weren’t good enough yet.

Deniz Kavi: We weren’t sure about what the applications would look like. What the– is the use case, production agents? Is it experimentation? Is it like proteins or small molecules or, how much of the scale part matter? And we had to invent that with our customers, and that was a pretty meaningful thing for us to go through over time. In that sense to invent a category basically. It was the most surprising experience for us was to figure that out. And unlike, if you’re starting a CRM company or like a pretty typical software company, you can just check what’s happening in the world and you do it– you could do it again, but better. In our case, it was very much a case of how does this product happen ever and how does it become actually applied? I think for the most part, there’s been a lot of efforts in this direction, but none of them have gotten that much actual usage.

Deniz Kavi: So we had to invent what the of that product would be.

Grant Belgard: What skills compound fastest at the biology software interface?

Deniz Kavi: I think I would put just like being an information sponge as a very valuable skill to have. I think people underappreciate how much– if you just know a lot about the use case you’re interested in, if you know a lot about the customer you get to or just what the product you’re building is. Oftentimes, again, the world is changing very quickly. You want to be on top of everything. My view has always been I’m gonna spend all the time I have in the world passively learning what I can and being as close, curious, and interested in people as much as I can be. This can be, whatever, social media posts, YouTube lectures, or, it could just be a scientist you want to talk to. think a benefit of being in the world is that you get to talk to experts all the time.

Deniz Kavi: And I think if you’re doing a computational science PhD, and all you do is like you re- you write papers and you just publish them and nothing happens afterwards. Talking– like having a person that uses and is depending on you is a very strong incentive to learn in general. And the ability to learn is a very powerful thing to have. And putting yourself in a structural position where you can learn a lot is just quite valuable. The PhD obviously like you’ll learn a lot there, but the things you learn are very much how do I invest– advance my general PhD research theme versus, people are relying on me that are quite important in their organizations.

Deniz Kavi: What can I do to just solve their problems or go another avenue on the learning side?

Grant Belgard: What’s worth learning deeply even though the tools keep changing?

Deniz Kavi: I think honestly, the– one of the strongest– one of the biggest gaps we have is like very strong core AI folks. I think you can see the incentives here, like if you’re a really good AI person, why would you want to work in biology? Basically, like you can do the core foundation model research, you can do some interesting robotics work or do computer vision, what have you.

Deniz Kavi: You’ll probably be paid more. There’s a couple of companies I think like bucking this trend. There’s a few like good research organizations, but I think the core foundational computational skills are quite important to learn right now. And there’s a lot of these getting replaced by the old LLM coding agents, but I think having that technical fluency is quite valuable in that, I think many bioinformaticians have become basically people in the interface of software and techno– software and biotech, is they become masters of none. And that is not a great place to be for most.

Deniz Kavi: You’re being bought– you’re being hired for a specific skill, or you’re starting a company for a specific skill, and if you’re okay at everything you might be a decent cog in a machine, and you might just be helpful in an organization, but you’re probably not gonna move the frontier of that field forward very much beyond that. I think founders become a lot of generalists. I think you do need to be a generalist. I fundraise and sell, and also you write papers about like our benchmarking results, what have you. And to some extent, being an expert in something that is to computer science and then applying that to life science problems can be a very valuable way to approach this.

Grant Belgard: How should early career people think about breadth versus specialization?

Deniz Kavi: I think my general concern about breadth, especially in our field, is that you are putting together a lot of things where I’ve taken some classes at Stanford that were basically surveys of the field because you like want to simultaneously appeal to computer science folks. You want to learn about bioinformatics.

Deniz Kavi: to apply t- to apply the computational scientists, learn about more computer science. And those courses are like useful, but they also become basically surveys or like very high-level walkthroughs of the field. And then those don’t tend to create a thoughtful scientist in and of themselves. I think my view in the world in general, like not even beyond this field, is if you want to be an expert in something, and then if you are an expert in one thing, you can then become an expert in another thing and go from there, as opposed to trying to do those things simultaneously. To some extent, I have some concerns about the… I have skepticism of inter-interdisciplinary degrees, for example, in that direction where like classes or majors that prefer to be doing these two, two things simultaneously.

Deniz Kavi: Because if you’re trying to learn two things at a time and then combine them together, the degree program like makes a theme or makes up a way to connect those together, but not necessarily a way to actually be a good credentialed expert for that use case. My view would basically be, become the smartest person you can be for one field. I would probably start with computer science, honestly, and then see where you can apply the science from there, is my general view.

Grant Belgard: How can people get close to real scientific pain points instead of building in the abstract?

Deniz Kavi: I think honestly, talk to people or put yourself in a place where you have to do it. I think in some ways one of the best ways to learn about the economics of a therapeutics company might be to start a therapeutics company. I don’t know that I would recommend doing that for everybody just to learn about that stuff. But I think I really believe in the power of being forced by your environment to do something hard and learning those things on the fly. I think if you’re reading papers about how the drug discovery, this process is done, like it’s just hard to be in that mindset. Like it’s even hard to motivate yourself to read those things, right?

Deniz Kavi: ‘Cause like it’s not gonna be immediately useful to you. So I think put yourself in an uncomfortable situation. I guess in the like general case, I have friends who have applied for jobs in places that they’re, they were definitely not qualified to do, and they offered geo service for free for them, and if somebody responded, they would, they have to make that work in the next one or two weeks, and they respond in that way. I think putting yourself in an uncomfortable spot like forces you to learn something, and like that might be a bit embarrassing also if you fail. this is how I think about those things.

Grant Belgard: How should scientists decide among various career paths?

Deniz Kavi: I think it’s just, it’s a question of, in my mind, I optimize on what I think is gonna be most impactful. The most people or has the most meaningful scientific sort of core that grows into something else, is how I think about that. Part of the reason I became a founder was very much in this context of if I did a research route, might be a good scientist like five or 10 years from now, but if I’m the median scientist, I’m not gonna be very important to the whole of science. And I’ll have a lot of papers with have 20, 20 citations, don’t really go that far was what I was concerned about personally. You can be a great scientist. I think academic research or, re- general research focus is very important, and some people will do very powerful, impactful jobs there work, work there, I should say. My view has always been what are you optimizing life? In my mind, that was impact.

Deniz Kavi: So however many people I can impact, I would like to optimize for that, was how I thought about that. So if your belief is that, the s-science in– of the next generation of, say virtual cells will be in academia, you should go all into that and then become a PhD and a professor and doing virtual cell work. If you think it’s gonna be applying some AI tools to therapeutics work, you can start a biotech company or start your own company to some extent. Or the other side might be, I think the most important bottleneck is clinical trials, even though I have this life science background, let me go try to find out what the clinical trials can be solved with. Yeah, I think so find the thing to optimize that you think is most valuable and go for that.

Grant Belgard: What mistakes do smart, ambitious people make when they enter this space?

Deniz Kavi: I think the… I think the mistakes smart people make are always almost the same in independent of the field. I think oftentimes the questions are you try to learn from somebody else, become a clone of what you think is a good thing to be doing. I think it’s a– for the most part, it’s like a decent way to live your life is to copy someone who is better than you and pretty close proximity to you. I think that’s reasonable, but to some extent, you do want to take a leap at some point. You do want to take a risk as much as possible, and that doesn’t necessarily have to be like, dropping out of school and starting a company, but it can also be something like, going beyond replicating somebody else’s work, I should say.

Deniz Kavi: I think we set ourselves like heroes and people to admire and work to replicate and to be similar to, but that is– that gets you only so far in what you can accomplish. It’s helpful to have knowledge of history and what things happened before, but you probably don’t want to be living in the shadow of somebody else or someone else’s work as much as you would in an empirical context.

Grant Belgard: What habits help you stay grounded when the hype cycle gets loud?

Deniz Kavi: I think I just might say talking to customers again. I think sometimes I’ll– people will tell me their negative results, so I think it’s quite go– There’s a lot of ex accounts who just talk about how biology will be solved next week. And like I am quite like– I think I’m very much exposed to that technology world when I’m quite excited by the technology in the first place. But once you know what’s happening in a field, it becomes much less exciting for the most part. I think, I suspect, people who do the core AI research for LLMs are probably much more conservative than the people making products on top of them. They’re probably both very excited and ambitious right now. I think we’re actually yeah, I think ke-keeping yourself in a place where, you know, people whose job it is not to sell AI for their company, but people whose job it is to apply it to their use case.

Deniz Kavi: This is very AI-specific, but also can go through most everything else is, when you open news-newspaper, you see the headlines about science, and you think they’re all bad, and you read all the, headlines about politics, and you think they’re all credible.

Deniz Kavi: I think it’s one of these things where you just want to be close to the source as much as possible if you care about the results from a given category.

Grant Belgard: What question do you wish more people asked you?

Deniz Kavi: I think a lot about why there have not been great software companies in the life sciences over the, say past 50, 20, 30 years. I guess to answer that, I think the value of software was somewhat limited. most of R&D spend in biopharma goes to development, like it’s, assays, experiments, headcount and then like a bunch more of that money goes to clinical trials and all the commercialization efforts for like obvious reasons, because it’s much more expensive. And I think a lot about like why we are different in this world where, basically our belief would be AI is sufficiently new that– or powerful that compute will be a more meaningful spend on the biotech side, and this company can be a whatever billion dollar revenue company in the next five, 10 years. And, I get asked this from like VCs a lot, but not so much by products and practicioners a lot.

Deniz Kavi: And I am curious how folks think about the tools they use, like how viable and sustainable the companies themselves will be. Like in many ways, a lot of these companies get bought by private equity, they get merged into some other thing, like the product does not improve over five, 10 years. And, if you’re scientists who use these tools, don’t think too much about what happens. It feels like whatever happens to the company happens to them also, and like they don’t really get to control what happens. And I’d be curious about how scientists think about what software tools they use exist as sustainable, long-living companies versus, a thing they bought for a year that might disappear next year.

Grant Belgard: What belief do you hold more strongly now than you did a few years ago?

Deniz Kavi: I think I’m a believer that AI will be valuable drug discovery. I think like this is a– I don’t see the starting company as a more open-ended question, because of the bottlenecks in clinical trials. Again is AI the right place or the– is drug discovery the right problem to apply this even in the life sciences world where– and what if you’re optimizing clinical trials, you’re like shaving one year off a clinical trial? I’ve become very convinced that at the current rate of getting better, the models improving, will have, net new assets added to every biopharma company’s portfolio of assets that will be, driving revenue for them. I think like we think a lot about, science outcomes like approvals and clinical trials, but I think, AI will make more money for pharma companies is on the drug discovery side by making more drugs for you to get to clinical trials and then approvals eventually.

Deniz Kavi: And I think this will almost inevitably will happen in the next of years or very soon, basically. I think, the adoption is there, there’s interest, but also I think the results will be speaking for themselves very soon.

Grant Belgard: What about the inverse? What have you changed your mind about in the last few years?

Deniz Kavi: Yeah, I think the– I am generally actually averse to science for like, goodness sake. I think you do want to think more about what the outcomes will be. Yeah, I think revenue is a core metric when you’re trying to make medicine, think about the commercial results. But to some extent you are serving tooling for an organization, and I have not thought about, I think about science as an interesting, valuable way to impact humanity and be good for the world, what have you, and I believe that still. But what you do day-to-day on– at your job is basically the same, whatever job you do. Like some– or between the roles I can do anyway are very much the same. And you want to pick your battles more carefully what you apply yourself to. I think I used to think that, you want in a field that is good for the world necessarily has to be strictly a very clear value add.

Deniz Kavi: if I went back to my, yeah, to three years ago and I start– just starting a company, I might just tell him, think more clearly about do you want to work in science forever?

Deniz Kavi: Do you want to make a bioinformatics software company? To some extent, you don’t necessarily want to be making a plan for yourself without knowing the requirements for that plan. So in this case, it would be, let me rephrase what I was trying to say, which is basically is the sort of aligning core thing you’re motivating your acts by?

Deniz Kavi: I mentioned previously you want to pick a job or a thing– or one singular thing that you’re optimizing for, but be more thoughtful about what that thing is and not just which fields are is good or perceived to be good by from a fuzzy, open-ended sense. And I think there’s a lot of value to be added for clinical trials, like a lot more boring even like payroll or like better health insurance, what have you.

Deniz Kavi: That type of thing is still quite valuable, and I would not– I would have looked down on those companies beforehand, where I’m now much more agnostic to humanity. This is a singular sort of effort to make the species better, and whatever you dare can be quite valuable.

Grant Belgard: What boring looking skill turns out to matter far more than people think?

Deniz Kavi: I think generally remembering things is quite important. I think like making structures for yourself to remember people’s names and where they work and what their, use cases are. I think it pays a lot of dividends in getting to know people effectively. If you keep seeing the same person over and over again, you can talk to them more effectively.

Deniz Kavi: Yeah, I think taking a lot of time to remember other people and also the specifics of their work or like what they might be related to you is quite valuable. We’ve seen this a lot in hiring, where like keeping track of somebody for many months, like just eventually converting them to a part of the team or a customer in our case.

Deniz Kavi: Or it might just be scientific collaboration or working with a partner on a specific program. I think just being very– taking interest in people and being thoughtful about keeping that interest in your– like in a structured way. Maybe you should remember from the top of your head, maybe you should write it down for yourself. I think paying attention, is the general theme here, is like purposely and spending a lot of time paying attention is quite valuable.

Grant Belgard: What are you still trying to get better at?

Deniz Kavi: Yeah, I think I spent a lot of time living and breathing with scientists, and I still do most of the time. I think we are as a company looking a lot more like how we can get this enterprise level adoption question of, a mature company in our space should look like, nine-figure revenue per customer. And that means basically everybody at this company uses you and like for a wide variety of core tasks. I want to become a better– I run a what, 15, 20 person company today. I’m not like a huge organizational figure. They’re all in the same room, basically. My– I want to learn about how our customers think about systems, how they think about the organizations of their large organiz- companies, like making those func- machines function well.

Deniz Kavi: Having managed a small number of people so far is a very complex problem, and both from understanding of what product we build next, I would like to be understanding of how the organization behaves and reacts to those things, but also running my own team more effectively. I think so I have basically over the course of the company, I’ve picked up skills I’ve never had and I was quite bad at the start, and I’ve become pretty good at.

Deniz Kavi: I think the next thing I want to be best at is understanding the behaviors and intents of a large organization and building products first and foremost for those organizations, but also apply those learnings to my own

Grant Belgard: What would be exciting for you to see happen in this space over the next few years?

Deniz Kavi: I think if we got 50% of the way to what the models claim they’re able to do, honestly, it would be very exciting. I think that would be a pretty meaningful proliferation of these models for different use cases. The approaching design is reasonably good, but if we got to I guess whoever those sort of hype people say on Twitter about these models, if they become real, which I think is like within the bounds of the we get today, would be very exciting. I think ultimately Tamarind is a index on the belief that, models will be getting better for molecular design applications, and it’s obviously good for me that the models get better, but also for the value of how we can produce more medicines more quickly, more effectively, and then even before that, just producing new medicines, period, so you can have another list item in the list.

Deniz Kavi: So I guess beyond the more boring and the models get better, but if they got as good as what people, some people think it, they are at right now would be the main thing I’m looking forward to.

Grant Belgard: Last but not least what would you like our listeners to take away from this conversation?

Deniz Kavi: Yeah, I think the general thing is you don’t really need permission to do things. You can start a company, you can work on a research project. That doesn’t mean that it’ll be easy. People will be rejecting you very frequently. But being in an environment where you are constantly rejected, being very– in an environment where you you’re punching above your weight is a very valuable place to be. I think people go and follow existing tracks a lot in life, which is very comfortable. It’s not pleasant, but what you need, what you’re supposed to be doing. I think you can go above your station and if you can hold yourself long enough there, you probably will ascend to a larger point. I am a not very long career– I did this out of school, and I started a pretty meaningful company still, and, much more to be done for sure. It’s a pretty, pretty small company still.

Deniz Kavi: And I think that the way– The limited success I have had came from being– putting myself in a place where I did not feel welcome or I felt like I was not supposed to be in this place.

Deniz Kavi: And oftentimes I was not supposed to be. Like, I was clearly not qualified for some of the things I began doing. And then just being in that situation makes You rise to the occasion when you’re put in an occasion where you have to be responding to something very hard, and putting yourself purposely in a place where you have to do something quite hard or at the time not possible for you to be doing is quite valuable.

Grant Belgard: Dennis, this has been fun. Thank you for joining us.

Deniz Kavi: Yeah, thank you for having me.