The Bioinformatics CRO Podcast
Episode 92 with Shannan Ho Sui

On The Bioinformatics CRO Podcast, we sit down with scientists to discuss interesting topics across biomedical research and to explore what made them who they are today.
You can listen on Spotify, Apple Podcasts, Amazon, YouTube, Pandora, and wherever you get your podcasts.
Shannan Ho Sui is a Principal Research Scientist in the Department of Biostatistics at the Harvard T.H. Chan School of Public Health and Director of the Harvard Chan Bioinformatics Core.
Transcript of Episode 92: Shannan Ho Sui
Disclaimer: Transcripts are automated and may contain errors.
Grant Belgard: Welcome to the Bioinformatics CRO Podcast. Today, I’m speaking with Shannan Ho Sui, principal research scientist in the Department of Biostatistics at the Harvard T.H. Chan School of Public Health and director of the Harvard Chan Bioinformatics Core. The core supports researchers through bioinformatics analysis, training, and platform development with a focus on high-throughput sequencing and reproducible collaborative research.
Grant Belgard: Shannan’s path spans biochemistry, genetics, iPSCs, pathogen genomics, cancer genomics, and now the leadership of a team working across many areas of modern computational biology. Today, we’ll talk about what a bioinformatics core does, how she built her career, and what advice she has for scientists working at the interface of biology, computation, and collaboration.
Grant Belgard: Shannan, welcome to the show.
Shannan Ho Sui: Thank you so much, Grant. It’s a pleasure to be here
Grant Belgard: So for listeners who’ve never worked with a bioinformatics core, how would you describe your current role?
Shannan Ho Sui: As the director of a bioinformatics core, we’re really there to support researchers with their data analysis. And so my role is to sit at that intersection between biology, statistics, and computing, and find solutions to their biological questions that they have from these large, high throughput, messy data sets. being at the core means that I have to put on a number of different hats at different times. As you already mentioned, we focus on consulting which involves analysis itself training, and also developing platforms. And so not only do I have to wear the different hats in terms of developing curricula, but also developing analysis strategies and also trying to think about reproducible ways to run pipelines.
Shannan Ho Sui: But there’s also that component of interacting with people and trying to really listen and understand what their biological question is, so that we can find the right methods that are appropriate for their data, while also managing sort of the finances of the core resource allocation. So a lot of different things that I have to do in my role.
Grant Belgard: What kinds of scientific conversations are you most excited to have right now?
Shannan Ho Sui: It’s an interesting question. I think there’s a number of different themes that have come up recently that have been very interesting. I’m, by training and at my heart, I’m a biologist, so the conversations I end up having are the ones that have a really interesting biological puzzle to solve, while at the same time, new technologies are emerging, and especially in the spatial and single-cell spheres.
Shannan Ho Sui: In the last few years, we’ve been doing a lot of work in that area, and it’s allowing us to answer these very interesting biological questions in new ways. And we have some great collaborations within the Harvard community. Scientists are doing really innovative, interesting work. Some of the work we’ve been doing recently is with Dr.
Shannan Ho Sui: Rachael Clark at Brigham and Women’s Hospital, who’s a dermatologist, but does a lot of research into inflammation and immune responses in skin. And so I’ve worked with her for, I think it’s 10 years now. But we’re still doing really interesting things looking at face transplant rejection and now being able to use spatial technologies and single cells to really characterize those pro-inflammatory and anti-inflammatory responses when you have an a transplant, and it translates to other transplant tissues as well. And also, I enjoy having conversations about given a specific biological question you’re trying to answer, what is the best technology to try to tackle this, and what is feasible and what is not? Because I think a lot of times researchers have great ideas, but it’s not always clear what the approach is to try to get at those questions.
Shannan Ho Sui: And sometimes what they think they want is not necessarily, the most appropriate. Trying to solve those puzzles is what I find most interesting.
Grant Belgard: What’s a misconception people often have about bioinformatics support?
Shannan Ho Sui: I think one of the things that we really strive to do in our core is to be at the level of a collaborator. So even though sometimes, you know, a bioinformatics core is classified as service, which it is also a collaboration. Very rarely I’ll be able to just run things or set an analysis pipeline and hand back the results.
Shannan Ho Sui: It’s very interactive. Um, we bring a lot of expertise both on how do we analyze that data, but also how to interpret it. so I think, I think that’s maybe a misconception. I think there are still folks out there who think, “Oh, I have a data set. I’ll just send it off. I’ll get some answers, and I’ll be off running, writing my paper.” especially with the complexity now that we’re getting from these new technologies. We plan to talk to our researchers a lot about what we’re seeing, how to interpret, we come, too, with a lot of biological knowledge. Within my group everybody has a biology background. And so there’s a real component of interpretation that’s key to supporting projects
Grant Belgard: What makes a project especially well-suited to a bioinformatics core?
Shannan Ho Sui: I think the best projects have a very clear biological question or hypothesis, even if we’re not completely sure how we’re going to analyze that data. It’s clear what we can define as success. The other thing that makes it well-suited for the core is if it’s something that we are already familiar with and something that we see frequently and have methods in place that are best practice. That’s not always the case, and sometimes we do projects where new technology is coming onto the scene and would benefit from the expertise that we can transfer over from other experiences we’ve had. but in those cases, it’s only suited for the Bioinformatics Core if the collaborator understands that there’s a component that is learning and development in addition to the analysis itself. And so I think that all translates into what I was saying before about it being a true collaboration.
Shannan Ho Sui: Need to be able to sit down, discuss the design, to iterate on results, and to invest time in communication, not just send data and get a figure
Grant Belgard: Related to that, what does an ideal first conversation with a collaborator sound like?
Shannan Ho Sui: I always start the conversations that I have with our collaborators with asking them what is your biological question?” Usually they come to me, and they already have a study in mind that they either are planning or are have already generated data for. And for me, the background on the biology is key to the whole thing being a success. So at the outset, I would like to know, what is the motivation for this study? How did you get to where you are now? And, what preliminary results or previous results led to you asking this question in the first place? And so then from there, I wanna be able to talk about… ideally, they would have not done the experiment yet, and they would be coming to me for a design consultation. And so then I will ask them, “Have you already thought about how you’re planning to answer this question?” that usually gets us into weeds of, what is the source material?
Shannan Ho Sui: How many replicates are they doing? What are the conditions? Do they have the appropriate controls in place? What are the key comparisons that they’re planning, to execute? And I think in that conversation also, just being frank about what we can do and what we can’t do, and where there are risks and where there are things that we think will be fairly straightforward. being able to talk about those things in a very open and candid way is important, and I think during that first conversation, if it’s an ideal conversation there will be a shared understanding that this is research, that bioinformatics is not going to solve all the problems.
Shannan Ho Sui: It’s going to help answer questions but we’re going to have to work together to get that done.
Grant Belgard: What are the earliest signs that a project could run into trouble?
Shannan Ho Sui: Oh, you know, just talking about this with my team recently because I think we’ve been re-scoping a couple projects, that went over budget, and one of the red flags I think is when… first I’ll say that estimating how long a project is going to take is incredibly difficult without knowing or having seen the data or the QC on the data.
Shannan Ho Sui: So I always tell people, “I don’t know exactly how long this is going to take. This is a difficult thing to do. We’re going to have to be willing to accept ranges, and we’re also going to have to accept that there are going to be these natural stopping points to check in, and we may need to reassess budget at that point, depending on what we’re seeing in the data.” If I get pushback on that or if it seems like the collaborator’s potentially inflexible on budget or inflexible on, know, how the analysis is going to be done that usually tells me that we’re likely to run into problems.
Grant Belgard: What you wish researchers would ask before generating data?
Shannan Ho Sui: I wish they would come to me to talk about the experimental design. Just having that detailed discussion on what they’re planning to do. I also sometimes … I’ve had people ask me this, “What do I need to do, as an experimentalist to make this collaboration go well?”
Shannan Ho Sui: Because there are things you can do, like being very well organized with your metadata is a key component of having a successful collaboration with the Bioinformatics Core. And so asking what format we need it in, that, that’s helpful. I think also, I always ask about timeline for the project because I want to know if there are upcoming deadlines. it would … It’s always helpful if they ask me too what is the turnaround time? And also h- let me know, this is not a question, but informing me if they expect to have long breaks in between. Because as you can imagine in academia, sometimes you do an experiment, you get it analyzed through the core, but then you may be preoccupied with other things that have come up and you’re juggling multiple priorities. analysts, it’s much more efficient for them if there’s a quick turnaround from their side, too.
Shannan Ho Sui: So having that discussion about timelines and what we can anticipate in terms of the communication and turnaround times is super helpful.
Grant Belgard: When, uh, they’re not submitting the paper three years later after lots of follow-up and coming back to you
Shannan Ho Sui: Yeah, no, that’s true. But also just, and I fully understand this. We have a lot of clinician scientists who we work with. When they tell me they’re going to be going on clinic for five weeks, I know I’m probably not hearing from them for that period of time, and that’s fine. just so long as we know that, okay, we’re, as a core we work fee for service, and it’s paid hourly.
Shannan Ho Sui: So interim period, I’ll be trying to find some small project or something well-defined that person can work on in the meantime
Grant Belgard: What kinds of questions are hardest to translate into an analysis plan?
Shannan Ho Sui: I think some of the more exploratory questions can be difficult to translate. Yeah so there are some things that are standard and best practice pipelines for analyzing data. So I’ll give an example of a spatial transcriptomics project that we worked on recently. it was a large study, I think something 20 to 30 slides. And it did have a hypothesis, and I’m not gonna I’m not gonna say exactly what the project was because I wanna protect our collaborators. But it was done more just to see what they might see out of the project rather than I have a clear hypothesis about a particular mechanism or a pathway that I’m interested in interrogating.
Shannan Ho Sui: And so while you can do all the standard things of QC-ing the data quantifying, doing differential expression, looking across time points or looking across slides Without knowing what they really want to focus on, and this was a brain study, so without knowing whether, it’s a particular region or a particular type of cell that they’re interested in, you’re going all over the map just looking for anything that the data can show you.
Shannan Ho Sui: And sometimes you’re lucky and there’s something very clear and interesting that pops out. But there are many times where there’s lots of different things that could be interesting or sometimes very little effect size, for example, in a study with a treatment. And so there it becomes a question of, okay now that we’ve gotten through to this point and we ha- we need to decide, what the hypothesis is and what we’re gonna focus on if there isn’t a clear objective, then the analysis plan becomes very difficult to continue with, so
Grant Belgard: What does successful collaboration look like from your side?
Shannan Ho Sui: One of the things that I’ve really valued in some of my collaborations is, number one, trust and mutual respect on both sides. Recognize that our collaborators come with, come to us with tons of amazing expertise, and they’re being incredibly innovative with their datasets. and if they’re able to be organized, like I was saying about the metadata and the data itself, then trust us to take that data and t- look at the quality and try to preserve as much of it as we can that is reasonable because as I mentioned, we worked a lot with clinicians.
Shannan Ho Sui: So sometimes, you don’t have a perfect experimental design, or you don’t have perfect samples. You’re working with what you have, and they are precious samples. I think if there’s trust that we’re going to try to get the most signal we can out of the data while also being rigorous and being objective about what really is, high quality enough to move forward with, I think that really helps the collaboration. And then the part that I’ve maybe value the most, which is then when we provide the results back, that there is a lot of discussion and interpretation on both sides, and that back and forth between this is, from the bioinformatician, this is what I’m observing. This is what I think is true signal, and then that being interpreted in its context.
Shannan Ho Sui: So for example, being like, “Oh, that’s a T cell signal, and I was expecting that because of X, Y, and Z.” Those, that I think if you can get the bioinformatician to understand the biology that they’re discovering through the collaboration and build domain expertise over time, which just makes them better and better, while at the same time providing the collaborator back with scientific insight into their dataset, I think that’s what I would consider a success.
Grant Belgard: What parts of the work are most visible to collaborators and what parts are most invisible?
Shannan Ho Sui: The most visible are all the visualizations and the figures, the nice pictures that we create to try to help them understand and also then for their papers. most invisible is all the challenges that you have when working with a bioinformatics dataset; often there’s a lot of data cleaning that has to happen. Sometimes I think it’s really obscured how difficult it can be to run an algorithm how long it takes to run a particular algorithm. I was just doing an InferCNV analysis on a single-cell dataset recently, it’s a large single-cell dataset. Tweaking those parameters to make sure that it’s the result is rigorous and reliable, each run was taking two to three hours.
Shannan Ho Sui: and sometimes it was running out of memory, and those are not things that, you necessarily share with your collaborator that, oh, I gave it 120 gigs, and that still wasn’t enough, and then I had to re-run again for three hours. So I think that component is not as visible, and I think the component that sometimes I wish there was a way to resolve this is that, within my team, we have a lot of internal discussions about the datasets we’re looking at.
Shannan Ho Sui: Yes, there is a primary analyst who’s been assigned to the dataset, but there’s almost always a few people involved in helping to look at that data and making sure that the interpretation makes sense or if there’s been a, an issue or a particularly sort of decision point that needed to be made other bioinformaticians weighed in on that.
Shannan Ho Sui: And I think by the time we send the report, it looks like everything was straightforward and easy, and it looks like it has these pretty pictures in it that make a lot of sense. But it probably took much longer than the collaborator thinks it did to get there.
Grant Belgard: What do you think about authorship credit and intellectual contribution in core supported work?
Shannan Ho Sui: I’ve been the director of this core now since 2015, and when I took it over, I think that co-authorship wasn’t as emphasized as it has become over the last 10 years or so. So when I became the director, I made a conscious effort to emphasize that, at the level that we’re collaborating and the amount of intellectual input we’re providing, that we are going to request co-authorship. And so over time I would say that for a majority of papers that get published where we’ve worked with collaborators, we are getting co-authorship. Obviously, that’s somewhere in the middle of the authorship list because we’re not the data generators. But I think it has become better understood that, without the bioinformatics expertise, a lot of these studies wouldn’t be published.
Shannan Ho Sui: And so it is a, a critical component of the collaboration. I’m not shy to ask anymore.
Grant Belgard: Yeah. What’s one thing listeners might be surprised to learn about running a bioinformatics core?
Shannan Ho Sui: I run my core like a small business. It’s like a startup. There are … All of our income or revenue is generated through cost recovery. So the amount of time we spend on a project is billed in hours, and the major resource or the major expense are people. And so being able to balance the books and to, to make that work, because as a NIH approved core, we have to break even.
Shannan Ho Sui: We’re a nonprofit, essentially. So we have to be within that 15% of break even to meet requirements. And so a lot of financial management. There’s a lot of project management, resource allocation. We do get involved in writing grants. And then there’s, that component of um, outreach and marketing and business development that I think that, you wouldn’t necessarily think about when thinking of a core facility.
Grant Belgard: What does reproducibility mean in day-to-day bioinformatics work?
Shannan Ho Sui: I think this is something that is really difficult and something that we’re always striving to improve upon. For me, reproducibility means that if you do a project and you’ve executed an analysis, that if somebody comes back several years later, that you can, one, you know where that project is and you can You still have copies of that data, but then you can run the pipeline that you ran on that data and get a very similar result. I’m not gonna say the exact same result, because tools change, versions change, and things like that. But in essence, that you can replicate or reproduce that analysis. for us, that means, good documentation in the code, committing the code and versioning it in GitHub, writing our code in things like Python notebooks or R Markdown, along with the interpretation so that it’s clear which figure was used to, come to a particular conclusion.
Shannan Ho Sui: yeah, having the data be safe somewhere and accessible, as I said before. We’re always working on this. I think it’s something that is overlooked a lot the time, and I also think it’s something that doesn’t come naturally to people, and you have to have processes and standard operating procedures in place for it.
Grant Belgard: Where do reproducibility problems usually enter an analysis?
Shannan Ho Sui: I think the long-running complex projects are the ones that are most difficult to reproduce especially when you have a rich data set, you might pursue many different
Shannan Ho Sui: analyses
Shannan Ho Sui: to answer a particular question and iterate on that, changing the parameters and things slightly, at which point it becomes really tricky to know, what you need to save and what you need to eliminate. And Making it clear what path was taken to get to the final result. But even in smaller project, I think reproducibility can be challenging. One of the things that the core struggles with is how long to keep data in a project. Theoretically, we’re not responsible for that data, the person who generated it is. And as you can imagine, as a core, we have a lot of people’s data, so we’re using up a lot of storage space. And so we do have to make decisions about when we let that data go, when we finally delete it.
Shannan Ho Sui: Deciding when a project is over a, is not straightforward at all because people can come back years later asking for a reanalysis or even just the data so that they can finally submit their paper. And so I think challenges can come really at any step there, at any step of, when you delete the data, where you put the code, which version of the code you, you keep as the, the final version. But even things like, for example, we use templates a lot in our group so that we can standardize on what we think is best practice for certain analyses. Sometimes there can be a problem just because somebody didn’t modify something sufficiently in the template to match the data because there was a copy and paste error or something. I think it’s just a really challenging thing, but something that we definitely strive to try to address in my group
Grant Belgard: How do you handle the tension between rapidly changing methods and the need for stable, reproducible results?
Shannan Ho Sui: That’s a tricky one. As a core we need to be able to standardize and but at the same time we also need to be able to keep up with emerging methods. So we have a lot of discussions internally about what best practice is, but we also work a lot with the community. And we are part of, for example, the NF Core community.
Shannan Ho Sui: We have been part of Bioconductor. We wanna know what other people are doing and to maintain what is the current state-of-the-art in analysis. But every time you change your pipeline or change your approach, it’s extra time and extra effort, not just to try a new tool or to benchmark it for an existing dataset, but also to then make that reproducible for the future. For the most part we are using, as I mentioned, our templates and what we have established as best practice, but we’re always on the lookout for what looks like it might be emerging as the next phase or the next best practice for an analysis. We’re fortunate to be in a community where
Shannan Ho Sui: there’s
Shannan Ho Sui: a lot of communication, there’s a lot of seminars, there’s a lot of discussion about what people could be or should be doing with datasets.
Shannan Ho Sui: And so I think you can never be complacent. You can never think that you’ve got it solved, but you do have to standardize and, … So I would say 80% is probably reproducible pipelines, and then 20% is pushing forward all the time to try to incorporate new things.
Grant Belgard: What should wet lab biologists understand before starting an omics experiment?
Shannan Ho Sui: I think one of the key things is before you do the omics experiment, is that what you need to answer your question? Some things could be solved with a qPCR or, with something much less expensive. And I think just there are, depending on the types of omics, being aware of things like batch effects and confounding and, Also being aware of, you can generate these really exciting datasets, but you also need to budget appropriate time analyze those datasets because I think a lot of times it’s very exciting to try a new technology. but if it is truly an emerging technology and has just come out on the market, there are going to be very few, standardized pipelines to analyze that data. And so then just being aware that they’re going to need to budget extra time so that somebody can really dig in and learn how to analyze that dataset well.
Grant Belgard: What bioinformatics concepts tend to unlock the most value for experimentalists?
Shannan Ho Sui: So the whole design, I think. So definitely understanding variation. When I do my consults, I will talk about replicates, but I’ll talk about that in the context of the variation that we might be expecting. So if they’re doing a clinical study, I know that there’s going to be a ton of biological variation.
Shannan Ho Sui: You’re going to need high numbers of replicates something from a cell line, for example. I think variation probably is of the highest value things that they can understand because even when you do time course for analysis there are different ways to set up a time course. You don’t always have to start at T zero for everything, right?
Shannan Ho Sui: You could end everything at the same time point, for example, and that might make sense in some cases. Yeah, I would say like technical variation and batch confounding, are high value because they can really make or break an experiment.
Grant Belgard: So we’re about to switch course for a bit and talk about trends in science and technology. But I just wanted to comment that so far your answers, if the tables were turned and you’re asking me the same questions you know, I would give very similar answers, probably not as eloquent, but essentially the same content, right?
Grant Belgard: I think everyone involved in providing bioinformatics as a service for long enough runs into the same issues time and time again.
Shannan Ho Sui: Yeah, no, it’s true. And I think yeah, I think if you’ve been in this field for a while, you would’ve seen how those decisions about the design really impact your ability to interpret that data. And yeah I would imagine that pretty much all core directors would be saying things very similar to me.
Grant Belgard: Which kinds of biological questions feel newly approachable because of current omics technologies?
Shannan Ho Sui: Yeah, so I think spatial is really giving us a lot of insight into things like cell-cell communication and local niches. And I think that’s really exciting. I think, single-cell was exciting because we had this ability then to get higher resolution into what might be happening, and we could look at signaling between cells, but it was still noisy because you couldn’t be sure that the cells were really in close proximity to each other or even in contact with each other. And so now I think that’s giving us the ability to ask some very well-defined questions. Again, I come back to this idea of if you have a really well-defined question, then that’s gonna help make the project a success. And so for example, if you have a hypothesis about, particular immune aggregates interacting with, maybe parasites or endothelial cells, you can now look at that.
Shannan Ho Sui: You can… and you can, subset your dataset to a particular niche and look very carefully at that. Whereas before we would, we could look at it at single-cell resolution, but it was all mixed in together. So I think we’re gonna make some very interesting findings using those approaches, and I’m excited about that.
Grant Belgard: How can scientists decide whether a more complex assay is actually worth doing?
Shannan Ho Sui: Yeah, so when we meet with people, that’s one of the key things that we’re always aware of. A lot of these studies are very expensive. so there’s always this balance of what is the question? What is the most straightforward way of getting to an answer for, to that question, and is it worth the cost?
Shannan Ho Sui: What are the potential pitfalls of pursuing a particular technology? Sometimes there are things that can create issues, like if you’re looking at a rare cell type and you’re taking a tissue section, how… i’ll ask very detailed questions like what proportion of those cells in the slide do you think you’re going to be able to detect? And how variable is that from slide to slide? Is that going to be in every slide, or y- do you just have to get lucky and get it just right to find what you’re looking for?” think, There are older techniques that not being used anymore that maybe could still be helpful instead of doing a very expensive spatial study. I haven’t had anybody doing laser capture dissection lately. I haven’t had any of those, but that was quite powerful for a while.
Grant Belgard: It is, yeah
Shannan Ho Sui: yeah. So I think it’s really thinking about the question and then finding the method rather than getting excited about a, a technology and… there is, as a bioinformatician, I’m always excited when somebody brings me something new to work on, but Weighing the cost of that with the, the return on that investment is important.
Grant Belgard: How can researchers make their data sets easier to interpret, reuse, and build on?
Shannan Ho Sui: Yes, I already mentioned the metadata. I think metadata is key. and lot of transparency. So I, there are the, the regular things that you can put in the datas, in the metadata, the types, the sex, the obviously. But then there’s all kinds of other information that the experimentalists can share.
Shannan Ho Sui: For example, the date of extraction for each sample, when were all the libraries prepared, even if it was done by two different people can have an impact. As much information as people are willing to provide, I’ll take all of it. So I think that is the key to the interpretation as well, because, the data will give you clues and tell you.
Shannan Ho Sui: We’ve had instances where we look at a study and we see something unusual in a couple samples, and when we go back to the researcher and we start to dig in and ask more questions, that’s when we find out, oh, for example, those two were the last ones that went onto the machine, right? Or onto the instrument.
Shannan Ho Sui: And so those, that sort of information upfront can help us so that we’re not spending time trying to figure out what the problem was, but we already know, okay, so they suspect that potentially these two samples might be different because they know they were last and there was like a, a bit of a longer break than they wanted between the first and the last samples.
Shannan Ho Sui: All of that aids interpretability. I think for single-cell data, domain expertise that people bring to the table are incredibly helpful. It’s getting easier to cell type with more datasets available in the public domain and also with
Shannan Ho Sui: AI.
Shannan Ho Sui: it still doesn’t beat domain expertise that biologists have built up over years. So telling us, we’re expecting these particular cell types, these are the markers that we trust most reliably for this, that helps a great deal.
Grant Belgard: How did you first find your way towards bioinformatics?
Shannan Ho Sui: I started in undergrad, I did my degree in biochemistry and molecular biology. And in my honors thesis rotation, was sequencing archaebacteria and looking for sites of RNA methylation. And I’m gonna date myself now, but, back then, a sort of bioinformatics involved running BLAST searches and sequence alignments.
Shannan Ho Sui: And I just found it incredibly interesting that you could identify what species you were working with from a couple of these searches using your computer, and it was very instantly gratifying. And around that time, I had a lot of friends who were in computer science. And so I was seeing them program, and I had started becoming interested in programming as well. And so then at that time, I actually thought I was gonna go to med school, and I was preparing sort of pre-med but decided at the last moment to take a second degree in computer science. And at that point I already knew that biology was what excited me and that I wanted to contribute to research. But it very quickly became apparent to me that I enjoyed programming and I enjoyed being able to answer some questions quickly, and that I wasn’t necessarily the best at lab work and going in on the weekend to check on my cells.
Shannan Ho Sui: and so I ended up moving into computational biology. So what had happened at that time is that my timing was fantastic because as I was wrapping up my second degree, that’s when–
Shannan Ho Sui: I’m from Canada, so I was at Simon Fraser University. But Simon Fraser and the University of British Columbia were putting together their, the very first cohort for a bioinformatics PhD program. And I happened to be designing a database for Fiona Brinkman, who was one of the people the PhD program, and she said, “You should definitely apply for this and move towards PhD in bioinformatics.”
Shannan Ho Sui: And so that’s what I did. And it was, know, best thing I ever did. I have so enjoyed it and have over that, this entire period excited about the work and the breadth of things that I can do, the variety.
Grant Belgard: Subsequently what were the key turning points in your career?
Shannan Ho Sui: Yeah. I did my PhD in genetics and bioinformatics. At that point, it was not even clear that bioinformatics would be a real discipline that you could do a, a PhD in. And my degree is actually in genetics, but it was all computational. I then had to decide what I wanted to do after the PhD, which was focused on gene regulatory networks and I ended up taking a position that involved a combination of research, but also project management.
Shannan Ho Sui: So I ended up being a, a scientist and a project manager for an initiative called Bioinformatics for Combating Infectious Diseases. And so that was my first foray into working in large collaborative teams and managing that. And and so that involved 11 different researchers in that consortium, and then having to understand what everybody was doing and corral people towards a common goal, was a great experience and it made it… know, it was with Fiona Brinkman. I actually went back to her after that to work in that position. It really showed me how, as a leader or a manager, you can leverage different e-expertise to do more than you could ever do on your own. And so after that, I ended up transitioning into Boston in working with Winston Hide in stem cell research.
Shannan Ho Sui: And so I was going from one domain to another domain, seeing all these different transferable skills, and then I had to make the decision of whether I wanted to, I think, have my own lab and pursue like a purely academic career or the opportunity came about that I could direct a core facility. And by that point, I’d already seen cores had access to such a wide variety of data and such a, an interesting set of questions that people were answering. And to be frank, I never had one particular interest in research that I wanted to pursue. And so I probably would’ve been a really bad professor because I didn’t have my own thing that I really was excited about.
Shannan Ho Sui: I was excited about what everybody else was doing. so that became the turning point where I essentially made the decision to direct the core and to support other people’s research rather than focusing on my own interests, which as I mentioned, I didn’t have a clear idea of what that was anyway.
Grant Belgard: Pivoting to issues in running a core, uh, what makes someone excellent in a collaborative bioinformatics role? What do you, what do you look for when you’re hiring?
Shannan Ho Sui: It’s a combination of a lot of different things. And I think it’s not just for cores, it’s for every and but it’s especially important in cores. So soft skills, I think, are incredibly important when you’re working in a core because you have to be able to really listen and understand what what people need for their projects.
Shannan Ho Sui: And so that involves, soft skills, deep biological knowledge, and technical ability. Those three things are incredibly important. I don’t think you can really be in a core without all three of those. And so I think maybe that’s also just the way that I’ve run the core that I want people to have that diversity in their repertoire.
Shannan Ho Sui: I know probably there are other models where you delegate specific things to people where their strengths are. For example, one person being more technical and then another person handling the communication. But for me, I think to be able to do all three of those things is incredibly important and valuable, and it just makes things much smoother.
Shannan Ho Sui: One of the things that I’ve also seen people excel is the ability to teach, not just communicate, but teach. And we have somebody on our team who is so articulate and eloquent at explaining things that it just makes the collaborations run very smoothly. And and I’d also say probably when it- in the realm of soft skills, the ability to take criticism and to also be able to stand your ground when you know that your approach is probably going to be the better one while at the same time being flexible.
Shannan Ho Sui: So I sound like I’m saying a lot of contradictory things, I think. But both flexibility and the ability to hold your ground, I think, are important if you’re going to collaborate.
Grant Belgard: What leadership lessons from bioinformatics cores would transfer well to biotech, pharma, or academic labs?
Shannan Ho Sui: Through my career, one of the things that I’ve found that people have found appealing from the experience I’ve gained as a core director the ability to build teams. So to understand what the needs are of the team and then to recruit people to fit those different roles. And so I’d say that that skill is obviously transferable to biotech to, to be able to resource small teams effectively so that they can be v- be very efficient and effective and successful.
Grant Belgard: Now the million-dollar question AI. How has it been impacting what you do? What are your expectations for where all this will go? What do you expect?
Shannan Ho Sui: Yeah. You and I had a brief conversation about that earlier. I think AI has a lot of potential benefits for bioinformatics. I think it’s great at helping people be better coders or to implement and execute code more quickly. Helps with the reproducibility, and there’s a lot of places where it can help with, you know, single-cell, for example, cell typing or anything that’s repetitive.
Shannan Ho Sui: I think AI or anything where you’re summarizing information, AI is incredibly helpful and a huge resource. It’s not particularly creative. So I think that in places where you’re looking for something a little bit more creative or more innovative AI is not going to replace that, although it can help make the more busy work or less interesting things faster. And as I mentioned in our previous conversation, one of the things that has come up is that while it’s speeding up certain components, it’s actually slowing us down in other ways. And a lot of our best practice for analysis has been put together through years of experience and through the community and talking to people and looking through t- at lots of different data sets with different nuances.
Shannan Ho Sui: But we’ve had occasions where collaborators take, the approach or the code, put it through Claude or ChatGPT, and then feel a sense of mistrust in the analysis because AI is suggesting something different. And I think that is something that we’re gonna have to learn to deal with and how to cope with that because, a lot of times AI will generate all possible solutions for a problem. And the one that we’ve selected is one that we likely have felt confident about, but now we have to explain why we didn’t choose all of these other ones. And that takes more time than I think that we often have and since we’re charging by the hour, it’s not really a good use of funds.
Shannan Ho Sui: So From that perspective, I think again, this is a communication and a p- and a people challenge, not necessarily an AI challenge. So teaching people how to use AI responsibly, which is something that we’ve been brainstorming coursework for recently. And we’re– we’ve already had a couple of courses on this, but we wanna flesh out that program a bit. But on the other hand, I think we’re going to have to use it. Like we… and I think that it has a lot of potential to make, our work easier and also to us better as bioinformaticians. So just really understanding it, understanding its limitations, understanding what it’s good for and where might want to be more critical is gonna be important.
Shannan Ho Sui: I often think, ’cause I, I have a young child what are the skills that we need to be teaching our young children for the future if AI is going to make many of the things we do now, easier and more automated? And I keep coming back to critical thinking. And for bioinformaticians, I think that means having deep biological knowledge because I don’t think you can think critically about biological problem without that.
Grant Belgard: Great answer. Shannan, thank you so much for joining us.
Shannan Ho Sui: Yeah, thank you so much for having me. It’s been a pleasure, Grant.







