The Bioinformatics CRO Podcast
Episode 89 with Beth Cimini

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.
You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.

Dr. Beth Cimini leads the Cimini Lab at the Broad Institute, where her group helps researchers turn images of cells into quantitative, reproducible biological measurements, and develops open source biological image analysis tools including CellProfiler, Piximi, and BiLayers.
Transcript of Episode 89: Beth Cimini
Disclaimer: Transcripts are automated and may contain errors.
Grant Belgard: Welcome to The Bioinformatics CRO Podcast. Today, I’m delighted to welcome Dr. Beth Cimini, who leads the Cimini Lab within the Imaging latform at the Broad Institute of MIT and Harvard. Beth works at the intersection of microscopy, computational biology, open-source software, and scientific community building. Her group helps researchers turn images of cells into quantitative reproducible biological measurements and develops open-source bioimage analysis tools, including CellProfiler, Piximi, and BiLayers. Beth’s career has spanned biochemistry, molecular biology, high-content imaging, and software, with a path from Boston University to a PhD at UCSF and then to the Broad.
Grant Belgard: We’ll talk about what bioimage analysis can teach us about biology, how scientific software and communities actually get built, and what advice she has for people who want to work across experimental and computational biology. Beth, welcome to the show.
Beth Cimini: Oh, thank you so much for having me. I’m delighted to be here.
Grant Belgard: For listeners who are new to bioimage analysis, what’s the core problem you spend your time trying to solve?
Beth Cimini: Yeah, so mostly we deal with images that come from light microscopes, and it feels like light microscopy should be a solved problem because we invented it in the 1600s, right? But it turns out that we invented light microscopes about 300 years before we invented computers, and so there’s a fantastic diversity of biology that we can study under the microscope. And cells might be anywhere from a pixel to a whole field of view. There’s tons of different fluorescent antibodies and can have many channels. And when you’re a bioimage analyst, you’re presented with an image and you’re sort of told, find the interesting biology here.
Beth Cimini: And so there’s a couple different pain points, one of which just being how can you find the objects of interest in the first place, which feels like it should be simple, but some of that’s because we have very good onboard neural networks, let’s call them, that help us find the boundaries of things, because otherwise we’d run into stuff all the time. Computer neural networks are still getting as good at object detection and object recognition as our brains have been, just because we’ve been doing that for a lot less time. And then there’s all sorts of sort of subtleties and nuances around, you know, how when you find the objects, can you measure them and interpret what interesting biology is happening. So it’s sort of a million dimensional problem, but it’s what makes bioimage analysis a lot of fun and very rewarding.
Grant Belgard: What does a typical week look like for you?
Beth Cimini: Oh, gosh. Well, so now I’m a PI, so now a typical week looks like a lot of emails and grants. But our lab is really fun. We’re sort of a four-legged table in that we do a mixture of open source software creation for image analysis, so things like CellProfiler and Piximi and BiLayers that you mentioned. We do image analysis methods research. So how can we make some of these things better and faster? We do image analysis almost as a CRO. So groups come to us, we collaborate with them from academia, pharma, biotech, wherever. And we do image analysis education and outreach. So a typical week, we’re touching all of those bits a little bit. And so we have an online educational platform that we’re building.
Beth Cimini: So it might be sort of checking out lessons that we’re building there, designing what the lessons are that we need to sort of put to make it easier for people to learn bioimage analysis, might be sort of checking in on a collaboration and seeing, you know, how’s it going for the workflow for, you know, these cool 3D images that, you know, our collaborator just sent over, and then sort of fixing software bugs and a little bit of everything. It’s chaotic, but it means it’s never boring.
Grant Belgard: What kinds of scientific questions tend to bring people to bioimage analysis?
Beth Cimini: Yeah, light microscopy is still one of the best things and certainly the most cost effective thing we have for doing anything where you want to do single cells. We’ve been looking at single cells under microscopes, again, for hundreds of years. So and especially anything where you want to do biology over time, we really can only of the omics that are out there only sort of reliably do sort of omics over time in light microscopy, because it’s the only omic we can do while the cells are still alive and still happy and still mostly doing their normal biology. But really, you know, we’ve estimated from some very back of the envelope things that we think probably about a third of biologists do some sort of light microscopy. So it can really be almost anything. Because we’re at the Broad and the Broad’s mission is to do things at scale.
Beth Cimini: We tend to work with people who are doing like early stage drug discovery, high content screens, but that’s a function of where we are. Biomage analysis definitely happens everywhere. It makes microscopy images hard to turn into trustworthy measurements. I think the thing that is tends to be the sort of most surprising to folks who don’t spend a lot of time thinking about it is, you know, for a lot of other modalities, we can very cleanly separate, you know, when we’re doing a sequencing read, and we see, you know, that we have a change, we have a nucleotide polymorphism that is sort of causing a biological change. Generally speaking, the change in the sequence doesn’t change our ability to do sequencing. Most biological changes won’t affect the actual measurement.
Beth Cimini: But when we get to bioimage analysis, we’re trying to look directly at the, you know, cell object or nuclear object or worm or whatever sort of object you might care about. And the phenotypic changes that are usually exactly what we want to detect, are going to affect our ability to find that object, potentially, you know, if we have a very finely tuned, you know, object detector or segmenter, to find, you know, cells, and we’ve told it, you know, cells are between 10 and 15 microns in diameter, if we have cool biology that causes the cells to be 17, our algorithm might just stop detecting the cell altogether. And so we will miss that cool biology, because it has gone outside the, you know, the definition of what a cell is.
Beth Cimini: And so it’s very hard to cleanly separate what’s a quality issue, and what’s an actual interesting biology, because we’re doing sort of detection of the things we care about, and sort of phenotype of the things we care about all in the same step. It’s not that we’re like, doing sequencing where we can measure if we’re getting a good read, and then sort of making inferences, you know, once we’ve collected our read, it’s the very things that we’re trying to detect as our biological, initial biological measurement, or the biology, there’s no intermediate steps. And so how can we do this well? And how can we make sure that we’re finding all the variability we want, and not other things like a, you know, 18 micron piece of crud that’s sitting there can be really challenging.
Grant Belgard: On that note, what makes a collaboration between a biologist and an image analyst successful?
Beth Cimini: That’s a great question. It’s one we spent a lot of time thinking about. It really, I think, just comes down to communication. We run office hours just for helping folks sort of, you know, get over particular small problems and things like that. But there’s a lot of nuances around, depending on exactly what kind of measurement you want to make. So example, if you’re taking measurements of colocalization, which are one of the most popular things to want to measure, you know, are two molecules in the same place, you know, maybe that means they’re interacting, maybe it doesn’t. There are certain kinds of controls that are really critical to do and will change the kinds of measurements that you can sort of reliably interpret if you’ve made them or not.
Beth Cimini: Which kinds of measurements are appropriate for which kinds of problems is sort of a thing that is like a passed down legacy from sort of bioimage analyst to bioimage analyst. But it’s not really anything where there’s hard and fast rules. There’s a lot of, you know, try this and, you know, under these circumstances, you know, this works and under the other circumstances. So it’s just a tremendous amount of, like, integrated knowledge that somebody who’s used to thinking about their biology as objects and not about how we turn those objects into quantitative measurements has just never contemplated. And so just lots of conversations. The good thing is all of the bioimage analysts that I know, and I know a lot of them are all, like, really, I describe us as a community of kind nerds who want to help people.
Beth Cimini: It’s really a field that tends to attract people who love collaborating and love talking about the sorts of stuff and sharing this knowledge and helping put people on the right path. It’s honestly the, like, the friendliest community I’ve ever been a part of. It’s really lovely.
Grant Belgard: How do you explain the difference between making beautiful images and extracting useful measurements?
Beth Cimini: I mean, they can be the same thing. They certainly can be. You know, a clean image that is mostly full of debris and stuff, you know, can’t, is probably going to be beautiful. Of course, beauty is in the eye of the beholder. But as a human, a lot of our, you know, visual system, we’re drawn to contrast. And so sort of people will do things like play with the gamma, which is, I’m doing hand gestures that, of course, the audience can’t see right now. Essentially, how much we change the value of the pixel that you see based on sort of the brightness of the photons that are hitting the detector, you know, because for human eyes, we like to see contrast, you might make it so that really small changes in pixel intensity sort of look very different and very striking and beautiful in an image.
Beth Cimini: But of course, we don’t want to artificially inflate the differences between two parts of our image when we’re trying to make detailed quantitative measurements. And so the things that will help you make quantitative measurements will also help make your images prettier in that, you know, they’re in focus, they’re clean of debris. But there’s a lot more we’re allowed to do if we’re just going for a pretty picture.
Grant Belgard: What kinds of projects are most fun?
Beth Cimini: Oh, gosh. I really like personally working with folks who are just getting into this space. So, you know, maybe this is the first time they’re sitting down and doing bioimage analysis, because just the like, the moment when it clicks and you see somebody be like, oh, I get it. I know what to do now. Like, that’s such a fun moment for me as a professional to sort of get to help somebody else get onto the path. I really enjoy that. But we’ve also been part of some some huge, you know, collaborative things like the JUMP-Cell Painting Consortium, which was organized by Anne Carpenter, also here at the Broad, where we were doing analysis for data made at 14 different sites.
Beth Cimini: And that was crazy to try to coordinate everything, but was also really cool to see, you know, it came out with sort of 200 terabytes of images and, you know, the biggest publicly available open source cell painting data set. And so things like that can be an awful lot of fun, too.
Grant Belgard: What kinds of projects are most dangerous to kind of, you know, be analyzed badly.
Beth Cimini: Oh, geez. That’s a good question. I think it’s really easy to do uncareful analysis, like, there’s a lot of big problems with bioimage analysis. There’s lots of hidden mine fields. I would say colocalization, I mentioned earlier, is one of the hardest ones. And that’s because there can be really subtle things with stuff like bleed through with stuff like, if I want to measure how often a little thing is inside a big thing, I need to take into account what fraction of the image is covered by little things and big things. So, you know, I have large spots and small spots. I want to know if the small spots are in the big spots, but I need to know if the whole image is big spots. They might be, but it might be by accident. It might be coincidental. And so colocalization is one of the things people most often want to do, but it’s one of the ones with the most traps in it.
Beth Cimini: So actually for this online educational resource that we’re doing, we started with saying we’re going to do a chapter about colocalization. And now we’re doing a whole course on colocalization because we know it’s really hard and we know there’s a lot of sort of subtleties to get right.
Grant Belgard: What makes cell painting useful for learning about cell state?
Beth Cimini: Yeah. So I can explain cell painting a little bit first for folks who’ve never come across it before. Cell painting is an assay developed here at the Broad in the early 2010s where the idea was, and it was, you know, more of a controversial idea at the time than it certainly is now. We have these cell images. They’re beautiful. They’re full of, you know, particular stains and particular markers to learn particular biology. But we certainly now understand from things like single cell sequencing and other just sort of computational domains, you know, just having high dimensional information, just lots and lots and lots of data means even if we don’t understand any particular specific data point, you know, we can draw connections between sort of samples or things like that. And so this is the way a lot of deep learning works.
Beth Cimini: We don’t really understand what a particular neuron in a particular deep learning architecture does, but we know when the whole thing comes together, you know, something interesting pops out. And so the idea was let’s put as many dyes as we can that will be inexpensive and fit on a standard microscope so that we can create really rich feature descriptions of cells. We’re not going to worry too much about what any individual feature means. And in fact, now in the era of deep learning, we sometimes use deep learning features that nobody has any idea what they mean. And just by doing lots and lots and lots of measurements on lots and lots and lots of cells will create big data big enough such that sort of interesting biology will fall out.
Beth Cimini: Now, I have to say, when I first came to the board, like I was trained as a sort of very standard cell molecular biologist and I was told about this assay and they were working on like the first one of the first big results papers from it. I was very skeptical. I kind of was like, there’s no way that works. Like, why did I just spend the last, you know, eight years of my PhD, like putting fluorescent proteins on particular things if all I had to do was dump a bunch of dyes on there and just sort of like study what pops out. But it turns out that, you know, making these dense feature representations is incredibly powerful. And so in that the eLife 2017 Rohban et al.
Beth Cimini: paper, which was one of the first big results papers from cell painting, sort of showing that you could detect novel pathway-pathway interactions between two different sort of gene pathways that were never previously hypothesized to interact just because they had a connection in cell painting feature space. I was like, whoa. And so now cell painting is a super well-adopted tech assay in the sort of early drug discovery space. It’s fun because it’s really sort of inexpensive to do. At scale, it’s less than $2 a sample, you know, compared to many thousands for a lot of other omics assays. It can be scaled up really easily. It’s amenable to every kind of mammalian cell we’ve ever thrown at it. And I know some folks who are working in non-mammalian cells who are using variants on it too. And you can get a lot of interesting biology out there. Now, it doesn’t work for every phenotype.
Beth Cimini: You can’t detect every kind of phenotype or every type of biology in it. But if you’re lucky enough that you can detect it in cell painting, then you have this sort of really great, really powerful assay that’s super easy to spin up and super easy to scale up if you decide you want to test it on hundreds or thousands of sort of drugs or genetic variants.
Grant Belgard: What kinds of biological variation are images especially good at capturing?
Beth Cimini: Yeah, I think there’s a lot that images are good at capturing. But, you know, a lot of stuff, because it has been the easiest modality to scale first, you know, a lot of what we know about biology comes from sequencing. And I work at The Broad, which is a place that grew out of the Human Genome Project. Sequencing is incredibly powerful and incredibly useful. But we know that there are things like, you know, regulation of translation. There’s an interesting paper that came out just a couple of weeks ago, Nature or Science, I forget which, about alternate translation. Alternate translation is actually a super common thing. Cell behaviors that happen on a fast state that happen faster than transcription or translation. You know, you phosphorylate a molecule, maybe it changes its behavior, maybe it goes somewhere else, maybe it gets sort of kinase gets turned on or off.
Beth Cimini: You can detect that with microscopy, you know, within seconds of it happening. For other omics, you have to wait for sort of like whatever the downstream activity is to sort of build up whether that’s, you know, you’re doing proteomics, and you’re sort of detecting, you know, new additions of ubiquitin or things like that. Those things take time. But for microscopy, you can do things sort of instantaneously, and you can capture cell-cell interactions, and you can capture cell behaviors in ways that are still really hard with other omics. The other omics are absolutely catching up, and all of the omics sort of working together is really our sort of vision of eventually where biology gets to.
Beth Cimini: But we are getting phenotype with microscopy at the level of like, okay, we can see how all the RNAs, all the proteins, all of the other things we mentioned in the other omics come together and make a cell behave.
Grant Belgard: What kinds of biological variation are images bad at capturing?
Beth Cimini: Yeah. So changes in sequence, obviously, are possibly detectable if they cause a downstream phenotype, but maybe not. I mean, it has to, for any sort of perturbation, we need to either have it change the cell enough that we can see a difference in something like a bright field image or a cell painting stain where we don’t know what specific changes we’re looking for, or we need to have a readout or a, you know, sort of something that’s designed to detect a sensor for a particular change. The good news is there are, you know, tons of sensors, and there’s a million antibodies in the RRID catalog, literally, but we need to know which one we want to detect. And unless you’re doing something like imaging mass cytometry or spectral microscopy, you can usually only detect four or five molecules at a time.
Beth Cimini: So if you want to see how 50 things are changing at once, that’s incredibly easy with sequencing. That’s incredibly difficult with microscopy. So parallelization is still, I would sort of sum up where I got to with that.
Grant Belgard: Where can image analysis create false confidence?
Beth Cimini: Oh, there’s a great paper on this recently from folks at Janelia called Believing is Seeing. So people assume because the computer did it, it is therefore unbiased. And having a computer analyze your images is definitely more unbiased than you sort of like doing what I was trained to do when I was sort of still a very young scientist, which was like, just like group all the pictures by like, by what they’re supposed to be pictures of, and then like decide what you think is the difference and then pick out some representative pictures and say representative image shown. That is the least, that is the least sort of unbiased you can be because a human is making those decisions. But again, things like I said about where you are making decisions about what the algorithms find.
Beth Cimini: It can be also really pernicious things like, well, I expected how, when do I decide my workflow isn’t working and decide to change it? Do I always do it no matter what? Or like, oh, in sort of run one of this assay and run two of this assay, I detected a threefold difference. And then the third time I run it, now I don’t see that difference anymore. Now maybe I decide to check, well, maybe my image analysis isn’t working anymore, but maybe it wasn’t working between one and two. And we just never checked it because the results were the same. And we were like, oh, it’s right. And so we only, you know, inspect really carefully when our results are unexpected. And so I definitely recommend reading that paper, seeing, believing is seeing, because if you believe it, you might find it. It’s a good way of how even with quantitative image analysis, you can trick yourself.
Beth Cimini: So it’s absolutely better to do quantification, but we’re humans, we’re biased. And, you know, we just have to, we have to design that into the systems that we build. And that’s something that we do try and design into the systems that we build to sort of make it easier to tell when the analysis is going wrong, but not in a way where we know the results and we expect what the results will be, but just in a more unbiased fashion.
Grant Belgard: What do you think about batch effects and imaging experiments?
Beth Cimini: Oh, batch effects and imaging experiments are like the thing that we are fighting against the most often, you know, especially with cell painting. The wonderful thing about cell painting is it is exquisitely sensitive. The horrible thing about cell painting is it is exquisitely sensitive. We’ve, you know, talking to many folks who are experts in this technique from all over the world over the years, you can tell with cell painting, which plates were in the back of the incubator versus the front of the incubator, which ones are on the incubator shelf versus like are sitting on top of another plate. And so you have all of these really exquisite differences that don’t matter, that are not what you want to detect, but they’re all mixed in with the differences you do want to detect, which are, again, the sort of things in biology.
Beth Cimini: And so being really careful with experimental design when you’re doing an imaging assay, especially something like cell painting, where you’re not always putting your controls in the same places or, you know, you’re making sure maybe you don’t even know what’s plated in each well, like you ask your friend to plate it for you and write it down. And you only sort of like you treat everything in a completely blinded fashion. It’s still really hard. And that’s why we do sometimes need to change image analysis algorithms sort of run to run or batch to batch, because there are batch differences.
Beth Cimini: And that’s, again, where that idea of like, well, is it am I now biasing the results every time that I change it, but I need to change it or it might be wrong next time, you know, you’re in this sort of like catch 22 of like, I want to keep it the same because I want my results to be comparable across batches. But also batches are different. So how can I sort of change it the exact right amount?
Grant Belgard: How do you approach rare phenotypes, subtle phenotypes, or like single cell heterogeneity?
Beth Cimini: Yeah, those are really hard. And those are things that largely as a field, we still haven’t, you know, come up with really good ways to deal with. I mean, one thing is always just kind of like, look at your data, look at your results, what level of results you can look at, you know, if you have millions upon millions of data points, you can’t look at them all one at a time, you can’t look at every image in something that has, you know, a million images, at least not without being there for three weeks. But figuring out how to sort of make sure things are, how you’re detecting the sort of rare things and how you’re making sure that you can even find them, that you’re not losing them in the segmentation process. Like I mentioned, you know, usually image analysis starts with, let me find objects that I care about, and your object detector might fail when weird biology is happening.
Beth Cimini: Hopefully, you know, something that sort of increases the likelihood of the weird biology, and you can go to look for it. But absolutely, I think we’re still only scratching the surface in terms of heterogeneity, in terms of how we deal with it, in the way that folks in like, the single cell sequencing field do in sort of large scale omics. At a smaller scale, certainly doing things like using super plots, where you can see all of the plots of your data, as opposed to like, not just like making a bar graph, bar graphs are the enemy. And sort of summarizations are the enemy can help with things like that. But it’s, it’s still something where I think we as a field need better methods. And I’m sure that, you know, smart people are working hard, and we will get better. But there’s still a lot of work to do there.
Grant Belgard: How can a team tell whether a model has learned biology rather than artifacts?
Beth Cimini: That is a great question. I wish I could answer for you. I mean, if there’s a particular piece of biology you want to, that you can induce that you know, you want to measure, then things are sort of like relatively straightforward, right? Because you sort of say, all right, I know that I should be able to induce, you know, a change in this marker, I have an assay that sort of like, you know, I have an antibody for that marker. And when I sort of add the drug or use the mutant that sort of where that marker should go up, it does. And again, doing things where you’re making sure that you’re careful to the person who designs the analysis is maybe blinded, you know, you’re not making sure that you grew all of your controls on, you know, last week, and you do all of your treatments this week, you know, basic experimental design stuff.
Beth Cimini: But especially when it comes to deep learning models, and especially when it comes to deep learning models, where we don’t have tons of data, you know, deep learning is going to learn exactly what it wants to learn. And your ability to tell what it learned can be really hard to figure out. There was a really great paper a couple years ago from somebody who had done a lot of work on cell painting stuff, and then use a tool called GradCam, they were sort of trying to classify cells with deep learning, and they use a tool called GradCam, which when you’re classifying a picture, it will sort of highlight in bright colors, the parts of the picture that it’s using to classify whatever you’re classifying. And it was saying the background is super important, not actually the cell.
Beth Cimini: And so in that case, they were able to check, they were able to sort of like, do this colored annotation of like important parts of the image. But when we teach, we have a bioimage analysis boot camp, we run a couple times a year. And we spend a couple of hours talking about how machine learning, you know, in general, and deep learning in particular, will lie to you, will learn what they want to learn, not what you want them to learn. And how can we out lazy the assay to try to figure out exactly what the model has learned? One of my favorite sort of like points to do is say, you know, what is the easiest way to train a classifier that you show it a picture of a cell and you ask it if the cell is interphase or mitotic? Just say interphase, because about 95% of cells, it turns out, are an interphase, and you’ll be 95% right. And most of us would take 95%, like 95% is a good day.
Beth Cimini: So if you just say interphase, and you don’t even interact with the thing at all, you’ll be right 95% of the time. And so trying to come up with counterfactuals, trying to sort of figure out what are ways the model could cheat, and design that into whatever algorithm you’re building, especially anything based off of machine learning are really critical, because it will cheat if you allow it to if it will find a way to do it. And we’re starting to see this now, I know, in some papers where people have taken LLM models, where it’s like, you give it an image, and it sort of tells you, like, this is a radiology image of such and such, and we see this condition in it. It turns out, if you don’t give them the image, they still say the same things. Like you you trick it into thinking it has the image, but it doesn’t, it will say the same exact things, whether it has the picture or not.
Beth Cimini: And so we need to be really careful to sort of make sure that we’re learning what we think we are.
Grant Belgard: Switching gears a bit, how do you balance tool building, tool maintenance, training, collaboration, and research? Beth Cimini Oh, gosh, I’ll tell you when I figure it out. I will say I have two fantastic team leads in my lab. Nodar Gogoberidze is our head of software engineering, and Erin Weisbart is our head of image analysis and training. And without them, you know, my life would be so much harder, and they sort of manage their respective domains really well. But yeah, there’s, and we have super smart, super hardworking people on the team, but it’s always a balance. You know, you want to be coming out with new features, but you also want to make sure that things don’t break. One of our software projects, BiLayers, is about sort of trying to promote containerization in software, because software containers are way less likely to break and are a way more portable artifact.
Grant Belgard: And so people might have to spend less time, you know, doing building, doing maintenance, because, you know, the things that they’ve made can be maintained more easily and for longer. It’s really a challenge to know, like, where in all of those things, like, am I spending my day doing the most good and helping the most people? So, so far, it’s just try a little bit of everything. But if you don’t build it, they won’t come. But if you do build it, and they don’t come, why did you build it? So you kind of need to be pushing on all of the pieces of the thing. And that’s why I like to describe it as a table. If we don’t have all of the legs of our table sort of in balance, the table isn’t a good table anymore. What makes scientific software sustainable? Oh, gosh, it largely isn’t. I will say scientific software sustainability is really hard.
Grant Belgard: And that’s a lot of it is because the vast majority of funding mechanisms are not aligned towards keeping things working. And I understand it. Like, if you’re like, well, would you like a shiny new thing? Or would you like to keep the thing that you have already working? You know, people want a shiny new thing. I totally get it. But at the same time, you know, you talk about a software project like the ImageJ project, which friends of ours make, which has about a million users a year and runs off of like a couple of people. And, you know, they’re always looking for ways to sort of fund that team better and be able to do more with that team. And this is something that a million people a year use. But nobody wants to fund something old when they could fund something new.
Grant Belgard: So I think that’s something that we as a sort of scientific community need to work on treating software as infrastructure. And we could fund infrastructure better in this country also and in the world also. But treating software as infrastructure and sort of saying in the same way that we know we need to maintain our roads and our pipes and our things in the digital age, software’s infrastructure too. And there should be more set asides to sort of keep stuff that way. You know, I mentioned containers earlier. That’s another way to at least make sure that things sort of can be used for longer. But new operating systems come out, new versions of languages come out.
Grant Belgard: And if existing tools don’t have funds, which funds equal time, then you end up with things like the heart bleed bug where, you know, open the SSL was being like done by one guy as a volunteer and something that half the internet relied on, you know, was insecure and nobody knew. So I think that’s something that we as a society could do better. And a couple of funding agencies are now starting to work in this space. But it would be great to have, you know, more focus on that being critical.
Grant Belgard: How can institutions build better career paths for bioimage analysts?
Beth Cimini: Yeah, I think first of all, they need to build career paths, not just better ones. Although I will say this seems like an area where the tide is turning. One of the first sort of big groups for bioimage analysis was the Network of European Union Bioimage Analysts or NEUBIAS, which was started in, I believe, 2015. And that was sort of one of the first times that people were like, hey, there is this career bioimage analysis. It exists. There are a few of us, like, let’s get together. And a few turned into a few hundred. There’s now a global bioimage analyst society, GloBIAS, that is trying to connect people from all over the world. But it’s a really strange career path in that, you know, it’s sort of, it’s by its definition, multidisciplinary.
Beth Cimini: You have people who started in biology or computer science or physics or math or, you know, something like that, and then pick up some of those other pieces to come to this, like, interdisciplinary space. And so it’s hard to sort of make a career path for it. But saying, hey, we have tons of our biologists making biology data, you should make sure it’s really quantitative and useful, is a thing that it now seems like more and more core facilities that I’m aware of are starting to have bioimage analysts on staff, which is, you know, a great step one. We ourselves have been running for a few years now a training program in bioimage analysis for postdocs, where we take awesome folks from the wet lab and train them in bioimage analysis. But I think definitely the sort of demand for help with quantifying images currently outstrips the supply.
Beth Cimini: But I think, you know, encouraging multidisciplinary rules, I think encouraging quantitative thinking in biology degrees, which having looked at lots of biology curricula from all over the country, you know, many biology PhDs still don’t require any statistics or computer science. And I’m not saying you need to be able to be a sort of like elite programmer when you leave, but being able to sort of have quantitative thinking be a really critical part of biology. I think there are a lot of fantastic biologists who got into biologists because it was the least math heavy science. And there’s fantastic biologists, and they do fantastic biology.
Beth Cimini: But I think as we go more and more into quantitative biology, saying, even if you’re not going to ever write a line of code, how can you structure your experiment to make it the most quantitatively interpretable is something a lot of people are never formally trained in, and they sort of pick up along the way. So I would love to see more biology programs sort of doing things like that.
Grant Belgard: What kinds of questions do people ask repeatedly? And what do those repetitions teach you?
Beth Cimini: Yeah, I think one of the most common questions that we get, so there’s an online image analysis help forum called image.sc, which if you ever, like, if you know nothing else about bioimage analysis, if you remember nothing else that I say today, remember image.sc. It is the central, like, help forum where you have more than 60 open source tools have one central help forum. And it’s a really friendly place to go and say, like, I have a picture, I want to know this about it, like, please, can somebody help me? And probably 10 people will. But we spend a lot of time, you know, sort of when we’re writing grants and things like that, chasing the hard problems, like the things that are currently really difficult to solve, you know, how can we do huge light sheet microscopy data that’s terabytes in size?
Beth Cimini: And one of the most common questions we get is just like, how with sort of really small data, can I determine what fraction of cells are positive for a particular marker, which is one of the easiest sort of image analyses to do, but people just sort of don’t even know where to start. And so I think we, it teaches me anyway, that like, we need to push the envelope on like the methods to solve things that are currently unsolvable, but that this education component of like, making sure that people know how to do the stuff that is computationally solved, but not necessarily solved for the person who needs to solve it and making it so that the tools are easy enough to use and the education is out there that people know the tools exist and know how to use them. We haven’t finished that part of the work.
Beth Cimini: And we can’t only focus on making the hot new tools because we need to make sure that the stuff that people are doing in their day-to-day lives, which is not necessarily the super hard stuff, is doable.
Grant Belgard: To talk about you, how did your own scientific interests change over time?
Beth Cimini: Yeah. For me, it really sort of began and ended with microscopy. When I was an undergraduate, sort of looking for a research lab, I was like, pretty sure I liked biology and wanted to do some research. I talked to a couple PIs and my undergrad PI, Bill Eldred, when I was interviewing with him, like pulled out a picture that they had taken of, I think this was a turtle retina, but like, you can just imagine this sort of like, brightly colored, lots of little specks all over the place, gorgeous image. And I was like, oh, that, that’s what I want to do. Like, I want to make those. And it started just as like, I wanted to do the microscopy and I still, I miss doing microscopy more than I miss anything else from the wet lab. But I got to graduate school and I wanted to answer a really fiddly, you know, thing. I studied telomeres, which are the caps on the ends of your chromosomes.
Beth Cimini: You have 46 chromosomes in each of your cells. And so you have 92 telomeres in each of your cells. And I wanted to say, you know, depending on how long the telomere is, so I have to measure that with one color of microscopy. There’s a protein that comes in two splice forms, two flavors, you know, does the length of the telomere change how much of protein flavor A versus B is there? So I have to measure three different colors very quantitatively in very small spots, a hundred of them per cell, hundreds of cells many times. And counting it by hand all of a sudden was not going to work anymore. So I had to learn to code and I really didn’t want to.
Beth Cimini: But once I did, I found the sort of like puzzle solving aspect of that to be so satisfying that I love doing bioimage analysis, I sort of picked up to the despair of my PhD advisor, all sorts of little side projects, like helping my friends analyze their data. And that when I found out when I was close to graduating my PhD, like this is a job, like I’m like, this is a job, like this thing that I love to do for fun. So I realized that collaborating with other people and helping sort of solve these image analysis puzzles, like was a huge joy to me. And that brought me to the Broad. And now I never dreamed of writing software, though, like our software that other people would use, I’d written some own code, you know, for some own software myself, I realized how much I personally enjoy and how much I personally get out of like making stuff that makes other people’s lives easier.
Beth Cimini: Like when I was a grad student, nobody was going to care at the end of the day, like what my project turned out, probably. But if I make tools that like make other people’s lives easier than every day, I’m excited to come to work, because I know that the work that I’m doing is going to lead to somebody else being able to make a cool discovery they wouldn’t have been able to before. And that for me is a way better reason to get out of bed. And a way better reason to come to work and work really hard.
Grant Belgard: What’s changed as your work has shifted from individual projects to leading people in programs?
Beth Cimini: I mean, it’s definitely now more of a, you know, you have to be thinking not just about how am I going to get through the next two or three weeks, but I have to think how am I going to get through the next two to three years and sort of really be planning for the long term because grant cycles are, you know, about a year long, you have to apply about a year before any money comes in, and you have a team of people and you want to make sure that you can keep the great people and pay them. You know, we can’t pay, especially our sort of computational folks, what they truly deserve, but we pay them as best we can with NIH budgets being what they are. And so really just sort of having this sort of like multi-level contingency plan of like, well, if I get this grant, then we’ll do this. But if we get this other grant, we’ll do that. And yeah, it becomes much more than just about you.
Beth Cimini: It’s about your whole team and how can you make sure that every single one of them has what they need. And so you can’t just be thinking about what we need now. We need to be thinking about what we need to do now to make sure we’re good a year from now, which is higher stress than just sort of thinking about what I need to do in the next three weeks. But the folks I get to work with are the best in the world. And I love getting to feel like I’m making their lives easier.
Grant Belgard: In your current role, which skills have you found matter more than you expected?
Beth Cimini: Multitasking for sure. And like fast task switching. The other thing, and it’s not a thing that I’m naturally very good at, is just documentation. Like leaving things in a place where if you need to task switch, you need to come back to something two or three months from now, like you actually remember where you were and what you did. And like, I was never the person who like loved writing in their lab notebook and sort of carefully documenting everything. But it became, you know, do this or fail, which was sort of how I came to coding too. So I guess I’ve had a lot of skills that it’s been like, well, you’re going to get better at this or you’re going to be bad at your job. So documentation was one that I really had to get much better at.
Beth Cimini: And I’m lucky to work with, I mentioned Erin Weisbart at my team already, but also Anne Carpenter are two of the best, like organized people who make the best documentation that I know. And they’ve taught me a lot as I’ve worked in this job.
Grant Belgard: Switching now to, to advice. What advice do you wish you had heard earlier? Or at least heeded earlier?
Beth Cimini: Yeah. I am grateful for all of the sort of like side projects that I got to do during my PhD. I wish I had realized earlier that the fact that I only enjoyed my side projects and not my main project meant that like a biology research career where I study a particular biological problem was not probably where my brain was happiest. And really not feeling like because that is the normal path, that is the path you must take. You know, I had very definitively decided I wasn’t going to be a PI and I came to the Broad to be a staff scientist. And then one thing led to another and ended up being a PI of a very different kind of lab, a lab where we do computational stuff. But I was, I had decided that these were the normal paths and therefore they’re the only paths available to me. And now I’ve been on this very unusual, strange own path.
Beth Cimini: And I’m happier than I think with any of the sort of normal paths, but it can be hard to imagine something besides what you’ve seen. So if you hate the things that you’ve seen, like go look for more stuff. There are weirder ways to get to a place where you’re happy than, than you possibly ever knew. You just have to find the people who’ve been on those strange paths.
Grant Belgard: I think I know the answer to this, but, you know, we always like to echo such things, right? So at what point in an imaging project should someone seek image analyst input?
Beth Cimini: Yes, early, as early as possible, preferably before you’ve done, once you’ve done your first pilot, and you should definitely be doing pilots, pilots with controls. We have a graphic in one of our papers. It’s a PLOS biology paper from 2023. The first authors are Senft and Diaz-Rohrer, where we show the circle of bioimaging and bioimage analysis. And they should be a circle. It should be that one thing feeds into the next, feeds into the next. And that your data, your image analysis and your data analysis helps you design the next best experiment. And that when you do your pilots, you analyze them all the way through to make sure that you can actually detect the statistical measurement you want to make. Otherwise, you don’t know if you can actually detect the thing you care about or not. Teams like mine have office hours.
Beth Cimini: A lot of places, your friendly imaging core will have a friendly local image analyst. If not, GloBIAS, the Global Bioimage Analyst Society, has a database of people who have agreed to be contacted on their website that you can just reach out and find a bioimage analyst near you who works on things that you work on. Because the sooner that you figure that out, the more likely that you’ll never end up sitting in a bioimage analyst’s office and then saying, the information you want just isn’t there. Those are the days I leave work the sad. It’s just when we have to tell somebody, I’m sorry, the information you want, the way the experiment was designed, we just can’t say anything.
Grant Belgard: So many parallels here with sequencing analysis.
Beth Cimini: Yeah. It’s so tricky. I hate having to do that, but sometimes the data just isn’t there. So the sooner you talk to us, the less likely that you’ll ever have that conversation.
Grant Belgard: What mistakes should people try hardest to avoid when collecting image data?
Beth Cimini: As we get more and more into people using automated microscopes, I think there will be less of this, but going in with a preconceived notion, you have to know a little bit. You have to know what stains and stuff should be present. But if you go in and you decide that you’re only going to take pictures in your controls of cells that look like this, and then you’re only going to take pictures in other fields that look like something else, you’ll find a difference. But it might not be the real underlying biology. So try to design things so that you can find more than just the preconceived idea you came in with. And it can be hard to do that while also making sure that you can find the thing that you know you care about, but it will help you sort of avoid some of those biases that will doom you before you can get started.
Beth Cimini: And that’s where I mentioned something like having your friend plate your samples for you so that you don’t actually know what’s in each well until after the experiment’s over.
Grant Belgard: How can early career scientists make invisible infrastructure work visible?
Beth Cimini: Oh, gosh, I think this is a thing that we as a sort of field still need to work on. But I mean, I think there are now more options for publishing things like the Journal of Open Source Software, like academic currency still runs on citations. And so making things citable, you know, make things like Zenodo DOIs, at least like put digital object identifiers on your work. So you can say, look at these things I made, they all have a digital object identifier. You know, some of it is, you know, administration has to meet us halfway. And you know, people who are hiring have to sort of, say, it’s not just about the papers, it’s about maybe what the papers do. But, you know, playing the game to sort of like get the metrics that you need.
Beth Cimini: Well, also, we as a community work to make it so that things like, you know, GitHub stars, or, you know, usages of tools like page visits to a blog that you wrote, that actually really helps people figure out how to do something like none of those are traditional academic currency, but all of them might be really valuable.
Grant Belgard: How should someone decide whether bioimage analysis could be a good career fit?
Beth Cimini: I think it’s really well suited to puzzle solvers. So if you love solving puzzles, I think bioimage analysis is a great career fit. I think it’s a relatively easy thing to just do, because there’s tons of images online in places like the Bioimage Archive or the Broad Bioimage Benchmark Collection. There’s tons of free tools, like, you can just sort of play and see if it feels like this is a fun thing for you. If you’re going to do it as a career, you probably have to be somebody who loves working with other people, if that’s the thing that excites you and not drains you. But I think if you like solving puzzles, and you like working with others, those to me are, and you’re organized and can sort of track having multiple things going on at once. Those for us are the things we hire for in bioimage analysts for helping them succeed.
Grant Belgard: Which bottleneck, if solved, would change bioimage analysis the fastest?
Beth Cimini: That’s a good question. It’s going to be a very, like, unexciting answer. But like, metadata standardization. I mentioned there’s like a million different things that could be in an image, and a cell could be one pixel, it could be the whole field of view. We have no common way to sort of describe what’s in an image, which means we have no common way to even say, like, should these two images look the same? Yes or no. And that goes to the thing about quality standards being hard. If we just had a common descriptive language, and smart people are working hard here. But if we all just agreed on how to describe the things, then, you know, having big databases of images and being able to make cool deep learning models and sort of suggest how your analysis should go would be so much easier.
Beth Cimini: But right now, you have to tell me and show me and we have to have a consultation and sit down and I love doing those. So I kind of don’t want to automate them away. But we could automate them away and make this a lot more approachable for everybody if we just agreed on a common way to describe things. It’s a very uncool answer, but it’s really what we need.
Grant Belgard: What worries you the most and what excites you the most about the next few years of AI in this field?
Beth Cimini: Yeah, I think I’m going to distinguish a little bit here between like deep learning in general and like generative AI, you know, large language models particularly. I think deep learning for bioimage analysis has proven incredibly powerful, especially in the context of segmentation. And people are doing fantastic work there. You know, we’ve got some papers out of like deep learning models that help you get more information about images. And all of that is great. Where I worry the most is large language models, at least the ones the common commercial ones that people use have been trained are designed to give you confident answers always. And bioimage analysis is a space that is just full of nuance. Very small changes in experimental design or in what you want to find make huge differences in workflows, make huge differences in the sort of measurements you might want to make.
Beth Cimini: And if you look at the where existing benchmarks are for bioimage analysis, and there aren’t many, LLMs are terrible at them. But because they’re advancing in other fields, because people like them in other fields, I worry about people confidently adopting tools that are not right for the job, without even realizing, doing their best, not saying like, well, this is wrong, but I’m going to do it anyway. But just saying, oh, well, the LLM says it’s confident that this is right. And I don’t think that the tools are there yet. And I think we need to understand it’s a very complex problem and not lose our skepticism. Because, you know, next probable word token predictor, you know, tells us that this is right. And those tools are powerful, people are using them to do cool things. But we need to recognize their limitations and not just sort of trust what they say.
Grant Belgard: What would you like the field to look like 10 years from now?
Beth Cimini: I want there to be slightly more professional bioimage analysts. I still think we need more of them because I think there’s still a lot of hard unsolved problems. But I do hope that a lot of the people who are currently stuck on the easy problems, which I mentioned, are actually like the majority of the problems, find it so that they can understand bioimage analysis themselves, they can learn it themselves, they can do it themselves, and then they can move on to something else. And maybe some of them will become the professional bioimage analysts who are working on the hard problems. And I think as microscopes continue to get better and better, we will come up with new hard bioimage analysis problems. But I hope that for the biologist who just wants to analyze the pictures and get on with their day, they can do that a lot more easily.
Beth Cimini: I think there’s, you know, again, we’re working on the table to make the tools better, to make the education better. And I hope over the next 10 years, I can look back and say, wow, things are a lot better than they were in 2026.
Grant Belgard: Finally, what should people remember from this conversation?
Beth Cimini: I hope what they remember is that microscopy and image analysis, like, are really powerful because they’re diverse. But therefore, it’s a discipline that’s full of a lot of like subtle pitfalls. And it’s good to talk to experts about it. But that the experts are out there, we exist, and we’re kind nerds who love talking about this stuff. So find us on image.sc, or find us in the GloBIAS Database, or come to my team’s office hours. We’re out there, we want to help you. And we really love talking about this stuff.
Grant Belgard: Thank you so much for joining us. It was a lovely conversation.
Beth Cimini: Yeah, thank you so much for having me. Have a wonderful rest of your day.





