Jefferson Parker

Desafíos y oportunidades en el análisis de secuencias durante el descubrimiento de fármacos

Una conversación con el Dr. Jefferson Parker, fundador de NullSet Informatics Solutions

A medida que la biología computacional avanza en el descubrimiento de fármacos, surgen constantemente nuevos retos y oportunidades. El análisis de secuencias ha sido durante mucho tiempo un aspecto clave de la bioinformática. En este artículo, hablamos con el Dr. Jefferson Parker, fundador de NullSet Informatics Solutions y experto en análisis de datos para las ciencias de la vida, sobre las nuevas fronteras del análisis de secuencias en el descubrimiento de fármacos.

CAS: Como biólogo computacional, ¿qué funciones ha desempeñado en el descubrimiento de fármacos?

Jefferson: He hecho de todo, desde dar soporte a laboratorios de descubrimiento de fármacos hasta trabajar con grupos de investigación traslacional. Más recientemente, trabajé con equipos clínicos en operaciones y desarrollo. He apoyado a equipos de farmacovigilancia con análisis de datos de seguridad, he trabajado en desarrollo de negocio y en inteligencia competitiva. Incluso he sido miembro integrado en equipos de ingeniería de software como analista bioinformático, una especie de traductor entre los científicos de laboratorio y los desarrolladores de software.

CAS: ¿Puede hablarnos del papel del análisis de secuencias en el descubrimiento de fármacos? ¿Por qué es importante?

Jefferson: En el descubrimiento, se puede utilizar el análisis de secuencias para el cribado de dianas terapéuticas. Al observar lo que ocurre a nivel transcriptómico, se puede volver a mapear cualquier conocimiento de rutas biológicas que se tenga y preguntarse: "¿Cuáles son las causas ascendentes más probables?". Esas son sus probables dianas terapéuticas o algo próximo a su diana en términos de señalización. Si su fármaco se diseñó para una diana concreta, con suerte le ayudará a confirmar que está dando en el blanco.

También puede empezar a buscar biomarcadores para la selección de pacientes. ¿Existe un perfil genético particular, ya sea a nivel de expresión o de secuencia de ADN, con diversas mutaciones? ¿Hay ciertas mutaciones presentes en los pacientes que los hacen más o menos susceptibles a que su fármaco sea eficaz? Esta es un área extremadamente emocionante y activa en la industria farmacéutica. Poder saber antes del tratamiento si un fármaco funcionará o no puede marcar realmente la diferencia entre el éxito y el fracaso. Además, no se hace perder un tiempo precioso a los pacientes. Especialmente en un campo como la oncología, donde los pacientes no tienen mucho tiempo, no se les hace perder el tiempo con ensayos y errores y múltiples líneas de terapia con fármacos que simplemente no van a funcionar.

Y todo eso implica análisis de secuencias. Afecta a todos los aspectos del proceso de desarrollo de fármacos.

CAS: ¿Dónde cree que están los mayores retos dentro del análisis de secuencias?

Jefferson: Diría que los retos están disminuyendo bastante rápido porque la tecnología avanza cada año. Antes teníamos lecturas muy cortas y el ensamblaje era un gran desafío. Ahora obtenemos lecturas más largas y, aunque el ensamblaje sigue siendo un reto, lo es menos.

Imagine que mete su ejemplar de Guerra y paz en una trituradora. Tiene fragmentos de página de un par de milímetros por un par de centímetros, por lo que es muy difícil volver a montar el libro. Pero si los fragmentos de las páginas fueran mucho más grandes, más largos, y tuviera un trozo de párrafo en lugar de un fragmento de palabra, es mucho más fácil recomponer la historia en el orden correcto. Eso es prácticamente lo que está ocurriendo ahora con la transición a lecturas cada vez más largas.

El almacenamiento sigue siendo un problema. Incluso en mi puesto más reciente, tuvimos que mover datos de secuencias, y la forma más rápida de hacerlo fue cargarlos en un disco duro y enviarlos por FedEx. En lugar de mover cientos de gigabytes o terabytes de datos por Internet, es más rápido enviarlos en una caja. El almacenamiento local no es el problema, sino la transmisión de una gran cantidad de datos de un lugar a otro. Hoy en día, una vez que los datos están donde deben estar, se puede encontrar suficiente potencia de cálculo para ejecutar el proyecto de secuenciación, pero hacer llegar los datos a las máquinas sigue siendo el cuello de botella.

Las muestras derivadas de pacientes también son un reto. Son dolorosas, las biopsias son invasivas y los enfermos no quieren tener que dar varias muestras. Una vez tomadas, generalmente se fijan en formalina y se incluyen en parafina, por lo que cualquier material de ácido nucleico estará degradado hasta cierto punto. Existen formas de intentar extraer y utilizar ese tipo de tejido preparado para la secuenciación, pero la calidad de la secuencia siempre se verá afectada.

Si usted es una empresa pequeña, la tecnología también es cara: estas máquinas cuestan mucho dinero. Del mismo modo, los biólogos computacionales son cada vez más comunes, pero siguen sin estar en todas partes, y todo el mundo quiere trabajar para los más grandes y mejores y ganar más dinero. La reserva de mano de obra está creciendo, pero sigue siendo limitada.

En cierto modo, ya nadie quiere trabajar en el análisis de secuencias. Todos quieren crear el próximo gran modelo de aprendizaje. El enfoque ya no está tanto en el procesamiento de números y el análisis de datos, sino en la IA y el aprendizaje automático avanzados. Todo el mundo quiere trabajar en la tecnología nueva, atractiva y brillante, y eso no es el análisis de secuencias. Así que eso será un reto pronto.

CAS: Do you really need a computational biologist to do sequence analysis now?

Jefferson: If you're doing cookie-cutter, well-established methodologies that are well-developed, validated, and documented, then no, you don't. You don't need someone who can carve a new wheel. There are plenty of off-the-shelf software solutions that can take input data from all the different sequencing machines. You load it in, drag and drop the icons for the pipeline that you want to process, and push go. You go and get coffee, lunch, or depending on the size of the machine you're working on, you might go home and come back in the morning, and it's done. You don't need somebody like me to do that. Any tech-savvy research associate can run it. 

If, on the other hand, you're working with a cutting-edge sequencing machine and trying to derive a new analysis methodology that has not been done before, then there is no off-the-shelf solution. There, you need someone who understands the biology–someone who understands the input data, the output data, the mathematics, and whatever else. For putting that all together and integrating it into a new solution that does not exist, then you need a “me” or someone who understands all those pieces.

CAS: You mentioned how people want to work with AI and ML now. Can these technologies be of help with sequence analysis?

Jefferson: With a well-curated data set, AI and ML can definitely help. I know for a fact there are organizations that are applying machine learning technologies to consume the literature and build out knowledge graphs, so that could definitely play a role. Could AI and machine learning help with just sequence assembly? Probably, but I don't know if that's really overkill. 

CAS: Speaking of AI, what are your thoughts on AlphaFold, which performs AI predictions of protein structures?

Jefferson: I think AlphaFold is an absolute game-changer. It gives you a much faster path to a structure, which can feed into things like computer-aided drug design much faster than you could ever do before. You no longer have to have an NMR or crystal structure to have a starting point. Is it as good as a crystal structure? Probably not. The real, measured thing is always going to be better than the simulation. But, in terms of time, you can have something available now. It's going to have an impact that we might not even be seeing yet. I feel like, with AlphaFold, the stone has been dropped into the pond, and it's made an impact, but these are only the first ripples that are forming.

CAS: What do you think is the cutting edge of AI and ML in drug discovery?

Jefferson: It's spatial, which is kind of the next generation of single cell. Multi omics. Bringing in the DNA, the RNA, the proteins, the metabolomics, and integrating all of it. Even integrating it with cellular pathways and intercellular communication. It's not just the single cell anymore. It's the single cell and the cells next to it and the cells next to those; how are they interacting? That's where it's headed, where it already is.

CAS: Do you think we're going to be creating models of biological systems?

Jefferson: If you'd asked me that when I was in grad school, I would have said humanity does not have the mathematics that can describe a biological system. Biology is complex chemistry, chemistry is complex physics, and physics is complex math. Everything is based on that. Physics is a mathematically solvable problem, it just requires an immense amount of data, and chemistry is, to an extent, the same thing. But biology… I used to believe that we did not, and would not, have the capacity to mathematically model a biological system.

But now, that is probably the direction it will have to go in. Will it require a quantum computer? Maybe? It might be after my lifetime, but I will say now with some confidence that at some point, humanity will be able to have accurate, reliable computational simulations of living systems. And that statement kind of scares me. I know there's a lot of work being done in the digital twin space. Limited first-step scenarios, but digital twins are online and being used in clinical trials now. That's kind of the beginning of it.

CAS: What do you think is needed to drive these new frontiers forward, then? Do we need new algorithms or new frameworks? Or is it really just trying to make everything fit together?

Jefferson: It's all of it—we need new ways of thinking about the problem. It may be applying old algorithms with new ways of designing or implementing new algorithms. For things like epigenomics and looking at DNA dynamics, or the non-coding RNA space, exome versus everything else, that's different from just sequence analysis. It's a different way of thinking about it. It's still the sequence, but it's not just the sequence. Those different ways of thinking about the issue will require different tools.

CAS: If you could wave a magic wand and solve one problem in sequence analysis and drug discovery, what would you solve? And what impact would that have?

Jefferson: I would make all the data well-annotated and available to everyone. All the proprietary data from companies, institutions, and universities… everywhere. On a well-annotated, well-documented, unified storage platform, freely usable by everyone. Because then there would be enough, and we could solve the big problems.

Jefferson began his research career at MIT, exploring xenobiotic metabolism in the gram-positive soil bacterium Rhodococcus aetherovorans. He got into computing when faced with an overload of data trying to annotate the genome to develop DNA microarrays, and he’s been working at the intersection of biology, computing, and mathematics since. His career has taken him through small pharma, large pharma, and consulting organizations, including Novartis and Thomson Reuters. Along the way Jefferson acquired his Graduate Certificate in Applied Statistics from Pennsylvania State University and a master’s degree in computer science from Boston University.

Now, Jefferson is forging a new path with his own bioinformatics consulting company, NullSet Informatics Solutions providing data and analytics, data modeling, and technology project management services.

Información relacionada de CAS

Aspectos destacados del seminario web de CAS Insights sobre el panorama de las patentes farmacéuticas

Glóbulos blancos atacando a una gran célula cancerosa rosada sobre una superficie de tejido rojo.

La nueva investigación traslacional en inmunooncología podría llegar a pacientes que no responden a las terapias actuales.

Cápsula farmacéutica que contiene una hélice de ADN naranja brillante sobre un fondo de placa de circuito oscuro.

Informe CAS Insights: desbloqueando el futuro de la innovación farmacéutica a través de la inteligencia de patentes.

Obtenga nuevas perspectivas para avanzar más rápido directamente en su bandeja de entrada.