NullSet Informatics Solutions의 설립자, 제퍼슨 파커 박사와의 대담
신약 개발 분야의 전산 생물학이 계속해서 발전함에 따라 새로운 도전과 기회가 끊임없이 나타나고 있습니다. 서열 분석은 오랫동안 생물정보학의 핵심적인 측면이었습니다. 본 기사에서는 NullSet Informatics Solutions의 설립자이자 생명과학 데이터 분석 전문가인 제퍼슨 파커 박사와 함께 신약 개발을 위한 서열 분석의 새로운 지평에 대해 논의합니다.
CAS: 전산 생물학자로서 신약 개발 과정에서 어떤 역할을 수행해 오셨나요?
제퍼슨: 신약 개발 연구소 지원부터 중개 연구 그룹에 이르기까지 다양한 업무를 수행했습니다. 가장 최근에는 임상 팀과 함께 운영 및 개발 업무를 담당했습니다. 약물 감시 팀의 안전성 데이터 분석을 지원했고, 사업 개발 및 경쟁 정보 분석 업무도 수행했습니다. 심지어 생물정보학 분석가로서 소프트웨어 엔지니어링 팀에 소속되어 습식 실험실 과학자와 소프트웨어 개발자 사이의 가교 역할을 하기도 했습니다.
CAS: 신약 개발에서 서열 분석의 역할은 무엇인가요? 왜 중요한지 설명해 주시겠습니까?
제퍼슨: 신약 개발 단계에서 서열 분석은 표적 발굴 스크리닝에 활용될 수 있습니다. 전사체 수준에서 일어나는 현상을 관찰하면, 기존의 경로 지식과 매핑하여 "가장 가능성 높은 상위 원인은 무엇인가?"라는 질문을 던질 수 있습니다. 이것이 바로 유력한 약물 표적이 되거나 신호 전달 체계상 표적과 인접한 지점이 됩니다. 특정 표적을 겨냥해 약물을 설계했다면, 서열 분석을 통해 해당 표적에 제대로 작용하고 있는지 확인하는 데 도움이 될 것입니다.
또한 환자 선별을 위한 바이오마커를 찾는 데도 활용할 수 있습니다. 발현 수준이나 DNA 서열 수준에서 다양한 돌연변이를 포함한 특정 유전적 프로필이 존재하는지 확인하는 것입니다. 환자에게 특정 돌연변이가 존재할 때 약물의 효과가 높아지거나 낮아지는지 파악하는 것은 제약 산업에서 매우 흥미롭고 활발하게 연구되는 분야입니다. 치료 전에 이 약물이 효과가 있을지 없을지를 미리 아는 것은 성공과 실패를 가르는 결정적인 차이가 될 수 있습니다. 또한 환자의 소중한 시간을 낭비하지 않게 해줍니다. 특히 환자에게 시간이 부족한 종양학 분야에서는 효과가 없는 약물로 시행착오를 겪거나 여러 차례 치료를 반복하며 시간을 허비하지 않도록 하는 것이 중요합니다.
이 모든 과정에 서열 분석이 관여하며, 신약 개발의 모든 측면에 영향을 미칩니다.
CAS: 서열 분석 분야에서 가장 큰 과제는 무엇이라고 생각하시나요?
제퍼슨: 기술이 매년 발전하고 있기 때문에 사실 과제들은 매우 빠르게 해결되고 있다고 봅니다. 예전에는 아주 짧은 리드(read)만 존재해서 조립이 매우 어려웠습니다. 지금은 더 긴 리드를 얻을 수 있게 되었고, 여전히 조립은 과제이지만 예전만큼 어렵지는 않습니다.
예를 들어, 전쟁과 평화 책을 파쇄기에 넣었다고 상상해 보세요. 몇 밀리미터 크기의 조각들만 남았다면 책을 다시 조립하기는 매우 어렵습니다. 하지만 조각들이 훨씬 크고 길어서 단어 단위가 아닌 문단 단위의 덩어리라면 이야기를 올바른 순서로 재구성하기가 훨씬 쉽습니다. 지금 리드가 점점 길어지는 추세가 바로 그런 상황입니다.
저장 문제는 여전히 남아 있습니다. 가장 최근에 근무할 때도 서열 데이터를 옮겨야 했는데, 가장 빠른 방법은 하드 드라이브에 담아 페덱스(FedEx) 상자에 넣어 보내는 것이었습니다. 수백 기가바이트나 테라바이트 단위의 데이터를 인터넷으로 전송하는 것보다 상자에 담아 배송하는 것이 더 빠릅니다. 로컬 저장 공간은 문제가 아니지만, 대량의 데이터를 한 곳에서 다른 곳으로 전송하는 것이 문제입니다. 요즘은 데이터를 필요한 곳에만 가져다 놓으면 서열 분석 프로젝트를 실행할 충분한 컴퓨팅 파워를 찾을 수 있지만, 데이터를 기기로 옮기는 과정이 여전히 병목 현상으로 작용합니다.
환자 유래 샘플도 어려운 문제입니다. 생검은 고통스럽고 침습적이며, 환자들은 여러 번 샘플을 제공하기를 꺼립니다. 샘플을 채취한 후에는 일반적으로 포르말린 고정 및 파라핀 포매(FFPE) 처리를 하기 때문에 핵산 물질이 어느 정도 손상될 수밖에 없습니다. 이러한 방식으로 준비된 조직 샘플에서 서열을 추출하고 활용하는 방법이 있지만, 서열의 품질은 항상 저하될 수밖에 없습니다.
소규모 기업의 경우 기술 비용도 문제입니다. 장비 가격이 매우 비싸기 때문입니다. 마찬가지로 전산 생물학자가 늘어나고는 있지만 여전히 어디에나 있는 것은 아니며, 모두가 가장 크고 좋은 기업에서 높은 연봉을 받으며 일하고 싶어 합니다. 인력 풀은 성장하고 있지만 여전히 제한적입니다.
어떤 면에서는 더 이상 아무도 서열 분석 업무를 하려 하지 않습니다. 모두가 차세대 위대한 학습 모델을 만들고 싶어 합니다. 초점이 단순한 수치 계산이나 데이터 분석이 아니라 고급 AI와 머신러닝으로 옮겨갔기 때문입니다. 모두가 새롭고 화려한 기술을 다루고 싶어 하며, 서열 분석은 더 이상 그 범주에 속하지 않습니다. 이것이 조만간 큰 과제가 될 것입니다.
CAS: Do you really need a computational biologist to do sequence analysis now?
Jefferson: If you're doing cookie-cutter, well-established methodologies that are well-developed, validated, and documented, then no, you don't. You don't need someone who can carve a new wheel. There are plenty of off-the-shelf software solutions that can take input data from all the different sequencing machines. You load it in, drag and drop the icons for the pipeline that you want to process, and push go. You go and get coffee, lunch, or depending on the size of the machine you're working on, you might go home and come back in the morning, and it's done. You don't need somebody like me to do that. Any tech-savvy research associate can run it.
If, on the other hand, you're working with a cutting-edge sequencing machine and trying to derive a new analysis methodology that has not been done before, then there is no off-the-shelf solution. There, you need someone who understands the biology–someone who understands the input data, the output data, the mathematics, and whatever else. For putting that all together and integrating it into a new solution that does not exist, then you need a “me” or someone who understands all those pieces.
CAS: You mentioned how people want to work with AI and ML now. Can these technologies be of help with sequence analysis?
Jefferson: With a well-curated data set, AI and ML can definitely help. I know for a fact there are organizations that are applying machine learning technologies to consume the literature and build out knowledge graphs, so that could definitely play a role. Could AI and machine learning help with just sequence assembly? Probably, but I don't know if that's really overkill.
CAS: Speaking of AI, what are your thoughts on AlphaFold, which performs AI predictions of protein structures?
Jefferson: I think AlphaFold is an absolute game-changer. It gives you a much faster path to a structure, which can feed into things like computer-aided drug design much faster than you could ever do before. You no longer have to have an NMR or crystal structure to have a starting point. Is it as good as a crystal structure? Probably not. The real, measured thing is always going to be better than the simulation. But, in terms of time, you can have something available now. It's going to have an impact that we might not even be seeing yet. I feel like, with AlphaFold, the stone has been dropped into the pond, and it's made an impact, but these are only the first ripples that are forming.
CAS: What do you think is the cutting edge of AI and ML in drug discovery?
Jefferson: It's spatial, which is kind of the next generation of single cell. Multi omics. Bringing in the DNA, the RNA, the proteins, the metabolomics, and integrating all of it. Even integrating it with cellular pathways and intercellular communication. It's not just the single cell anymore. It's the single cell and the cells next to it and the cells next to those; how are they interacting? That's where it's headed, where it already is.
CAS: Do you think we're going to be creating models of biological systems?
Jefferson: If you'd asked me that when I was in grad school, I would have said humanity does not have the mathematics that can describe a biological system. Biology is complex chemistry, chemistry is complex physics, and physics is complex math. Everything is based on that. Physics is a mathematically solvable problem, it just requires an immense amount of data, and chemistry is, to an extent, the same thing. But biology… I used to believe that we did not, and would not, have the capacity to mathematically model a biological system.
But now, that is probably the direction it will have to go in. Will it require a quantum computer? Maybe? It might be after my lifetime, but I will say now with some confidence that at some point, humanity will be able to have accurate, reliable computational simulations of living systems. And that statement kind of scares me. I know there's a lot of work being done in the digital twin space. Limited first-step scenarios, but digital twins are online and being used in clinical trials now. That's kind of the beginning of it.
CAS: What do you think is needed to drive these new frontiers forward, then? Do we need new algorithms or new frameworks? Or is it really just trying to make everything fit together?
Jefferson: It's all of it—we need new ways of thinking about the problem. It may be applying old algorithms with new ways of designing or implementing new algorithms. For things like epigenomics and looking at DNA dynamics, or the non-coding RNA space, exome versus everything else, that's different from just sequence analysis. It's a different way of thinking about it. It's still the sequence, but it's not just the sequence. Those different ways of thinking about the issue will require different tools.
CAS: If you could wave a magic wand and solve one problem in sequence analysis and drug discovery, what would you solve? And what impact would that have?
Jefferson: I would make all the data well-annotated and available to everyone. All the proprietary data from companies, institutions, and universities… everywhere. On a well-annotated, well-documented, unified storage platform, freely usable by everyone. Because then there would be enough, and we could solve the big problems.
Jefferson began his research career at MIT, exploring xenobiotic metabolism in the gram-positive soil bacterium Rhodococcus aetherovorans. He got into computing when faced with an overload of data trying to annotate the genome to develop DNA microarrays, and he’s been working at the intersection of biology, computing, and mathematics since. His career has taken him through small pharma, large pharma, and consulting organizations, including Novartis and Thomson Reuters. Along the way Jefferson acquired his Graduate Certificate in Applied Statistics from Pennsylvania State University and a master’s degree in computer science from Boston University.
Now, Jefferson is forging a new path with his own bioinformatics consulting company, NullSet Informatics Solutions providing data and analytics, data modeling, and technology project management services.




