Uma conversa com Nicole Stobart, Jeff Wilson e Mark Schmidt da CAS
Estruturas de autoridade e identificadores são utilizados pela CAS no setor químico há quase 100 anos como a base da nossa curadoria e indexação, líderes mundiais. Nossa equipe de ciências biológicas está agora buscando criar as mesmas ferramentas de indexação para um novo setor. Neste artigo, conversamos com Nicole Stobart, Gerente Sênior de Capacidades em Ciências Biológicas; Jeff Wilson, Ph.D., Gerente Sênior de Ciência de Dados; e Mark Schmidt, Líder de Governança de Dados, sobre como eles estão trilhando esse novo caminho usando estruturas de autoridade.
CAS: Como vocês descreveriam as estruturas de autoridade?
Jeff: As estruturas de autoridade, da nossa perspectiva, são fortemente focadas na identidade. Dentro de qualquer domínio, seja de proteínas, substâncias químicas ou ácidos nucleicos, você quer ser capaz de identificar entidades de forma única. Historicamente, em nossa coleção química, temos o CAS REGISTRY®, que é nossa autoridade clássica. Dentro do CAS REGISTRY, queremos descrever produtos químicos diferentes de forma única e garantir que, cada vez que encontrarmos a mesma coisa em nossa curadoria, ela seja identificada dessa maneira. Um exemplo simples seria o paracetamol. Quer você o chame de paracetamol, Tylenol ou acetaminofeno, todos são identificados como a mesma substância química.
CAS: Por que as estruturas de autoridade são tão importantes para os cientistas da vida?
Nicole: Nossas estruturas de autoridade atuais funcionam bem para pequenas moléculas, mas o que acontece com proteínas ou enzimas? Sim, existem sequências que podem ser associadas a elas, mas será que uma sequência com uma alteração de aminoácido é uma entidade diferente? Reconhecemos que os pesquisadores estão enfrentando dificuldades com isso. Existem diferentes empresas e outras organizações que tentaram organizar isso, mas ninguém fez uma organização totalmente autoritativa desses biológicos. Quando tentávamos levar informações biológicas aos clientes, descobrimos que não era possível sem organizá-las de forma autoritativa primeiro. Não se trata apenas de coletar pilhas e pilhas de dados, mas também de organizá-los de forma autoritativa e garantir que a maneira como você os organizou seja a forma como o resto do mundo gostaria de vê-los organizados.
Mark: No espaço das ciências biológicas, estamos tentando identificar todas as coisas importantes e chegar a um consenso sobre o que elas são e como chamá-las — isso é, na verdade, uma autoridade. Em nosso negócio tradicional, se eu tenho uma substância química, posso consultar nossa coleção sobre ela e adicionar novas informações. Estamos tentando viabilizar o mesmo para as ciências biológicas, com coisas como proteínas e enzimas, e para fazer isso, precisamos primeiro concordar sobre as identidades delas.
Jeff: Depende do pesquisador e de onde ele está, mas acho que podemos ser muito mais consistentes e confiáveis nos dados que apresentamos se pudermos descrever as coisas de forma inequívoca. No nível mais simples, voltamos ao que sempre fizemos: organizar dados para busca, de modo que, quando você acessa nossas ferramentas de pesquisa, como o CAS SciFinder, não importa se você quer chamar de câncer ou sarcoma, nós conhecemos a relação entre esses termos. Não forçamos você a pesquisar com todos esses termos diferentes para encontrar os dados; nós cuidamos disso em segundo plano, e você pode usar apenas um dos termos para encontrar tudo.
Além da busca e recuperação, quando você entra em implementações mais sofisticadas, como grafos de conhecimento e a descoberta de relacionamentos importantes, é difícil construir um grafo de conhecimento preciso se suas entidades estiverem espalhadas por vários nós. Se tenho uma proteína que me interessa como alvo, quero que todos os dados dessa proteína sejam armazenados em um único nó e que outras entidades relacionadas fiquem em nós individuais. Caso contrário, você obtém um grafo de conhecimento complexo e ineficiente, onde você tem 15 nós para essa proteína, 12 para aquela doença e 7 para a mesma substância. Você cria esse grafo complexo e não consegue perceber que, basicamente, existem 3 entidades ali, relacionadas de forma previsível, porque essas identidades não foram definidas previamente.
CAS: Como vocês definiriam curadoria na CAS?
Mark: As pessoas tendem a usar o termo curadoria para significar duas coisas diferentes. Primeiro, há a ingestão e agregação de dados, o que geralmente não chamamos de curadoria aqui na CAS. Depois, há o tipo de curadoria com a qual a CAS tem um longo histórico, onde cientistas analisam fontes originais de informação e aplicam esforço intelectual para extrair e melhorar essas informações, disponibilizando-as de uma forma mais estruturada.
Jeff: A maioria dos nossos cientistas está envolvida na curadoria que Mark descreveu, onde extraem informações importantes de fontes primárias e, em seguida, melhoram, padronizam e conectam esses dados a outros pontos, mas há uma camada além disso. Temos um grupo separado de pessoas que faz a curadoria das próprias coleções de autoridade, analisando novas informações à medida que chegam e decidindo: isso é algo novo? Não é? Como organizamos isso? Como elas se relacionam? Que terminologia usamos? Não temos as pessoas que extraem as informações da literatura primária tomando essas decisões; em vez disso, temos pessoas que fazem a curadoria da coleção de autoridade.
Mark: O ponto principal é que a autoridade nunca está concluída. Você está sempre adicionando coisas novas que foram descobertas recentemente ou que se tornaram interessantes, e também está adicionando novos sinônimos para mantê-la útil.
CAS: Qual vocês diriam que é a vantagem das estruturas de autoridade para a curadoria?
Jeff: Depois que você começa a agregar identidades e terminologia com estruturas de autoridade, fica mais simples identificar novas instâncias. Assim que você coleta todos os nomes para sua entidade, pode criar uma estrutura de autoridade para visualizar e combinar instantaneamente. Sinônimos facilitam a curadoria.
Mark: As estruturas de autoridade ajudam a organizar e agregar todas as informações em torno delas. Portanto, no caso das proteínas, talvez não chamemos essa proteína exatamente como você, mas se soubermos como você a chama e como nós a chamamos, podemos compilar todas as informações sob a mesma identidade — isso nos permite reunir todas as informações.
CAS: Can AI be leveraged to help create authority constructs or curate with them?
Nicole: We need to make sure that we have authoritatively organized and identified whatever it is that we're talking about to enable AI curation. It's really difficult to rely on any sort of machine learning or machine-curated data if it has not all been manually organized or at least thought of and identified in an authoritative way.
Jeff: We have a general philosophy about AI and how it relates to our data in that we leverage the most advanced technology we can to enhance human capabilities. We have hundreds of scientists here who are doing curation work, and if you look at what they do today, it's a lot more sophisticated than what they did 30 years ago. Each time we adopt a new technology, we use that to enable humans to do more sophisticated work. We leverage technology in natural language processing and machine learning to help identify documents and insights, but all of that is then presented to humans, who make the final decision about what's important or what's relevant and how it gets organized.
Once we’ve created that structured data, we do the same thing with technology and AI solutions on the other end. We try to leverage the best technology to show you the things you've asked for and things that are related to it. We can use predictive technology to help you plan syntheses, we have AI engines that help predict what prior art would be, and you can do Boolean-type searches and try to find things. So we're taking our highly structured data, and layering AI technology on top of that—not using AI to replace people.
CAS: How is CAS going to approach defining and identifying biological entities?
Mark: So this is where we start to talk about identity semantics. It’s a really difficult and complicated problem in life sciences, and we are completely aware of that, we are dealing with that on a case-by-case basis every day. We might not get this perfect, but we're going to do our best. We're going to make reasonable, defensible choices, which are useful to the people in the life sciences space. We will not please all of them every time, but we're going to be predictable and clear about it, so people can use the system that we deliver reliably.
When we get to questions like, “Is a one nucleotide difference a new identity or not?”, that’s a choice to make. But if three different single nucleotide polymorphisms become one identity instead of three, we absolutely need to describe all three of those differences and attach them to the one identity that we present. So even when we decide that different things fall under one identity, all the information is still going to be discoverable, connected, and accessible.
Initially, we will choose our identity semantics and define them, then as new information arrives, we will decide, “Is this a new identity, or does this add to an existing identity?” We want to utilize as much computing power as we can, but at all times, our ultimate source of truth for those decisions is going to be humans who know the subject area best. That’s how we are going to deliver a trustworthy service.
Jeff: You want to create a clear, consistent, and comprehensive rule set upfront. For people who haven't tried to define entities and aggregate information, it feels very arbitrary, but we find that when you're organizing information, you want to err on the side of being clear and consistent. You can find nuances of the science in specific cases, but implementing rules for everything causes things to get scattered, and makes things harder to find and collect. Instead, it’s better to create a rule that works for 99% of everything. In the chemistry space, we sometimes get correspondence from a scientist who says, “I see you described this thing this way in CAS REGISTRY, and you're wrong, I have data that says it’s better described a different way,” and our usual answer is that our content is organized around discoverability. And that’s the point, the curation needs to consistently lead you to the data you’re looking for, even if it misses some scientific nuances along the way.
CAS: Are you using structure or function to create these identities for life sciences?
Mark: For many biological entities, especially proteins, it’s more function than structure, but it’s often a combination of both. In chemistry, we can easily create authority constructs based entirely on structure, but that’s not the same for life sciences. We have several different authorities we have to solve for—proteins, antibodies, organisms, etc. Each of them will need to be approached in their own unique way. We have to move away from just structure-based identities as we move into life sciences in a way that we have not had to in the past.
CAS: How do you plan to tackle the longevity of defining by function when biological function changes over time and with new research?
Jeff: We always try to future-proof things, and I don’t think we can predict where life sciences will be in ten years, but, as we create consistent constructs, we’ll build some flexibility and extensibility into that. The key to this will be recognizing when to keep using the current construct and when a new branch of science emerges that requires a new construct. People won’t stop wanting to describe proteins, but there might be a subcategory of some kind that arises that needs new authority constructs to be created.
Mark: I think that if you do a solid job of getting reasonable identity semantics at the beginning, then you set yourself up for a future where it's most likely that scientists will subdivide one of your identities into a few more specific identities. If you think about genus-species naming of organisms, it worked for a really long time. The need for sub-species didn't completely invalidate the genus-species model. I think we can set constructs up where they might get more specific about identity than when we started out, but they won’t become completely outdated.
CAS: What makes CAS the right organization to aggregate these fragmented life science authority constructs?
Jeff: There is some self-sustaining nature once you become a worldwide authority on something. If you look at chemical substances, nobody second-guesses CAS REGISTRY. It's the authoritative place. There are some other substance collections, but we are positioned as an organization in a way that most other organizations are not. We are part of the American Chemical Society, our mission isn't just to be profitable, our mission is to facilitate science. We have the people, the expertise, and the space to tackle this life science data and organize it in the way we did with CAS REGISTRY. I think anything short of us being the authoritative space for proteins when this project is complete is us falling short. That's what we're going to do.
Mark: The desire for agreement on a meaningful set of identities within life sciences is pretty universal, and it’s a problem we’ve already solved in small molecule chemistry. Life scientists are looking for a clear set of identities and the relationships between them to be defined so they can organize data around that structure. We see ourselves wading into that complexity and bringing order to the chaos. When we get far enough with that, people will accept that we know all the proteins that matter to most people and talk about them in a way that most people can make use of. When they don’t see their protein in our collection, they’ll bring it to our attention, so we can add it rather than starting their own entirely separate database. That’s what we’re aiming for—not just having an authority construct collection but being an authority within the life sciences space.
CAS: How is what you’re doing different than what’s already been done?
Mark: I think, in a lot of scientific spaces, people solve a problem for themselves and their colleagues, locally, in a way that works for them. They create a database that has the identities of things they care about, described in a way that they understand, and the information about those. Meanwhile, someone in a slightly overlapping area who's doing some of the same work, and some different work, is choosing different ways to talk about those identities and assembling different information. Then we start to see projects unite some of those smaller spaces to get all those people on the same page and facilitate discovery across those boundaries. So that's already happening in life sciences, but there hasn't been an approach to bring a large amount of it together with one uniting set of identities.
The goal is to get to a place where you can come in with your protein target, with no awareness that it is a member of three or four biological pathways, but when you've found that identity within this set of information, you now see all those connections. Additionally, you might also learn that it's being used as a biomarker for a disease state or treatment outcome. Being able to aggregate all the information from different places around one identity, which you can then find and use, creates value that wasn’t previously accessible.
Jeff: Existem muitas pequenas organizações por aí que agregam uma parte do quebra-cabeça. A ontologia de doenças existe, o Uniprot tenta organizar proteínas e o NCBI possui um banco de dados de organismos. Cada um deles é uma autoridade por si só, mas não estão conectados de forma útil — você precisa ir de um lado para o outro. Estamos buscando um conjunto de dados unificado onde você possa encontrar todas as informações das ciências biológicas conectadas e harmonizadas. Para isso, precisamos pegar a visão do Uniprot sobre a proteína, a visão do Ensemble e a visão do CAS REGISTRY, e harmonizá-las em uma única visão e um único conjunto de identidades. A partir daí, você teria acesso a todas essas coleções menores. O resultado é maior que a soma das partes quando estão conectadas, pois revela coisas que antes não podiam ser encontradas.
Nicole: Queremos ser o banco de dados de ciências biológicas mais abrangente do mundo. Queremos resolver quaisquer problemas e desafios que nossos clientes enfrentam, e eles enfrentam essa necessidade de estruturas de autoridade nas ciências biológicas.
CAS: Vocês tiveram alguma resistência de outros bancos de dados existentes ao entrar neste espaço?
Mark: Não dá para agradar todo mundo. Em algum momento, é preciso exercer autoridade e fazer escolhas. É difícil discordar de algo bem estabelecido e consolidado, e tentaremos não fazer isso ao longo do caminho. Não vamos convencer todos a mudar a forma como chamam as coisas, mas queremos construir a partir disso.
O princípio é simples: vamos nos alinhar aos bancos de dados existentes. Mas, na prática, é definitivamente mais difícil fazer essas escolhas. Se dois bancos de dados usam termos diferentes para a mesma coisa, podemos escolher um ou outro, ou optar por algo novo e tentar melhorar ambos. Sei que parece muito difícil, mas sentimos que a dificuldade em chegar lá é exatamente o valor que vamos agregar ao setor e o que queremos entregar aos clientes.
Nicole: Definitivamente não queremos alienar as pessoas que usam bancos de dados existentes e encontram valor neles. O que queremos é agregar mais valor ao que já existe. Ainda não conversamos sobre como outros curadores pensam a respeito, então como eles se sentirão ainda é uma incógnita!
Jeff: Na verdade, a minoria dos cientistas se dedica à curadoria, então, de certa forma, eles não são as pessoas com quem precisamos nos preocupar, já que buscamos ajudar a maioria. E o que ouvimos de muitos cientistas é que eles têm dificuldade em acessar os dados de que precisam.
Nicole: Com certeza! Recentemente, conversei com um cientista que criou um banco de dados de informações sobre anticorpos. Ele disse que fez isso não porque queria organizar dados de anticorpos, mas porque precisava deles para executar seus modelos preditivos. Acho que essa é a situação de muitos cientistas; é um obstáculo para a pesquisa, e é isso que pretendemos corrigir.
CAS: Se você pudesse usar uma varinha mágica para corrigir uma coisa sobre estruturas de autoridade, o que seria e qual impacto isso teria?
Mark: O problema que eu resolveria seria fazer com que as pessoas usassem as mesmas palavras para significar a mesma coisa. Se chegássemos lá, tudo seria muito mais fácil. Grande parte deste exercício consiste em pegar os termos que os cientistas usam para descrever algo, encontrar a identidade correta e conectar os termos a essa identidade. Se pudéssemos padronizar a linguagem e fazer com que todos concordassem, poderíamos pular essa etapa.
Jeff: Para mim, seria fazer com que mais pessoas entendessem as estruturas de autoridade e tivessem a visão e a paixão para criá-las de formas úteis. Mesmo dentro de uma organização que apoia isso, ainda passo muito tempo explicando por que são necessárias e qual é o valor. Embora seja gratificante defender isso, acaba desviando a atenção da minha parte favorita, que é trabalhar com dados e construir coisas.






