与来自 CAS 的 Nicole Stobart、Jeff Wilson 和 Mark Schmidt 的对话
近百年来,CAS 一直在化学领域使用权威结构和标识符,将其作为我们世界领先的标引和索引工作的基石。我们的生命科学团队目前正致力于为这一新兴行业创建同样的索引工具。在本文中,我们与生命科学高级能力经理 Nicole Stobart、数据科学高级经理 Jeff Wilson 博士以及首席数据管理员 Mark Schmidt 进行了交流,探讨他们如何利用权威结构开辟这一新路径。
CAS:您如何描述权威结构?
Jeff:从我们的角度来看,权威结构非常注重身份识别。在任何领域,无论是蛋白质、化学物质还是核酸,您都需要能够唯一地识别实体。在我们的化学品收藏中,我们一直使用 CAS REGISTRY®,这是我们的经典权威系统。在 CAS REGISTRY 中,我们希望对不同的化学品进行唯一描述,并确保每次在标引过程中遇到相同事物时,都能以同样的方式进行识别。一个简单的例子是对乙酰氨基酚。无论您称其为对乙酰氨基酚、泰诺还是扑热息痛,它们都被识别为同一种化学物质。
CAS:为什么权威结构对生命科学家如此重要?
Nicole:我们目前的权威结构对于小分子来说运行良好,但蛋白质或酶的情况又如何呢? 是的,可以将序列与它们关联起来,但仅有一个氨基酸差异的序列是不同的实体吗?我们认识到研究人员正为此感到困扰。虽然有不同的公司和其他组织尝试对其进行整理,但还没有人对这些生物制品进行过完全权威的组织。当我们试图向客户提供生物信息时,我们发现如果不先进行权威性的组织,这是不可能实现的。这不仅仅是收集海量数据,更重要的是进行权威性的组织,并确保这种组织方式符合全球通用的标准。
Mark:在生命科学领域,我们试图识别所有重要的事物,并就它们是什么以及如何称呼它们达成共识——这才是真正的权威。在我们的传统业务中,如果我有一种化学物质,我可以查看我们的收藏并添加关于该物质的新信息。我们正试图在生命科学领域(如蛋白质和酶)实现同样的目标,而要做到这一点,我们首先需要就它们的身份达成共识。
Jeff:这取决于研究人员及其所处环境,但我认为如果我们能明确地描述事物,我们所呈现的数据将会更加一致和可靠。 从最简单的层面来看,这回归到了我们一直以来的工作:为搜索整理数据。这样当您使用我们的搜索工具(如 CAS SciFinder)时,无论您想搜索“癌症”还是“肉瘤”,我们都知道这些术语之间的关系。我们不会强迫您使用所有不同的术语来查找数据,我们在后台处理这些关系,您只需使用其中一个术语即可找到所有相关信息。
除了搜索和检索之外,一旦进入知识图谱和发现重要关系等更复杂的应用场景,如果您的实体分散在各个节点上,就很难构建准确的知识图谱。如果我将某种蛋白质作为研究目标,我希望该蛋白质的所有数据都存储在同一个节点上,而其他相关实体则位于各自的节点上。否则,您会得到一个极其复杂的知识图谱,其中同一个蛋白质有 15 个节点,同一种疾病有 12 个节点,同一种物质有 7 个节点。您创建了一个复杂的图谱,却无法看出本质上只有 3 个实体,且它们以可预测的方式相关联,因为这些身份在最初并未被定义。
CAS:您如何定义 CAS 的标引工作?
Mark:人们倾向于用“标引”来表达两种不同的含义。 首先是数据的摄取和聚合,在 CAS,我们通常不将其称为标引。其次是 CAS 长期以来所从事的那种标引,即科学家查看原始信息来源,投入智力劳动来提取和改进这些信息,并以更结构化的方式提供出来。
Jeff:我们的大多数科学家都从事 Mark 所描述的标引工作,即从原始来源中提取关键信息,然后进行改进、标准化并将其与其他数据点连接起来,但在此之上还有另一个层面。我们有一个专门的团队负责维护权威收藏本身,他们会查看新涌入的信息并进行判断:这是新事物吗?不是吗?我们该如何组织它?它们之间有何关联?我们使用什么术语?我们不会让那些从原始文献中提取信息的人员来做这些决定,而是由专门负责维护权威收藏的人员来完成。
Mark:关键在于,权威工作永无止境。 您总是在不断添加新发现或具有新价值的事物,同时也会添加新的同义词以保持其有效性。
CAS:您认为权威结构对标引工作有什么优势?
Jeff:当您开始利用权威结构聚合身份和术语后,识别新实例的过程就会变得简单。 一旦您收集了实体的所有名称,就可以创建一个权威结构来进行即时查看和匹配。同义词极大地促进了标引工作。
Mark:权威结构有助于您组织和聚合围绕它们的所有信息。 因此,以蛋白质为例,我们对某种蛋白质的称呼可能与您不同,但如果我们知道您如何称呼它以及我们如何称呼它,我们就可以将所有信息汇集在同一个身份下——这使我们能够将所有相关信息整合在一起。
CAS: Can AI be leveraged to help create authority constructs or curate with them?
Nicole: We need to make sure that we have authoritatively organized and identified whatever it is that we're talking about to enable AI curation. It's really difficult to rely on any sort of machine learning or machine-curated data if it has not all been manually organized or at least thought of and identified in an authoritative way.
Jeff: We have a general philosophy about AI and how it relates to our data in that we leverage the most advanced technology we can to enhance human capabilities. We have hundreds of scientists here who are doing curation work, and if you look at what they do today, it's a lot more sophisticated than what they did 30 years ago. Each time we adopt a new technology, we use that to enable humans to do more sophisticated work. We leverage technology in natural language processing and machine learning to help identify documents and insights, but all of that is then presented to humans, who make the final decision about what's important or what's relevant and how it gets organized.
Once we’ve created that structured data, we do the same thing with technology and AI solutions on the other end. We try to leverage the best technology to show you the things you've asked for and things that are related to it. We can use predictive technology to help you plan syntheses, we have AI engines that help predict what prior art would be, and you can do Boolean-type searches and try to find things. So we're taking our highly structured data, and layering AI technology on top of that—not using AI to replace people.
CAS: How is CAS going to approach defining and identifying biological entities?
Mark: So this is where we start to talk about identity semantics. It’s a really difficult and complicated problem in life sciences, and we are completely aware of that, we are dealing with that on a case-by-case basis every day. We might not get this perfect, but we're going to do our best. We're going to make reasonable, defensible choices, which are useful to the people in the life sciences space. We will not please all of them every time, but we're going to be predictable and clear about it, so people can use the system that we deliver reliably.
When we get to questions like, “Is a one nucleotide difference a new identity or not?”, that’s a choice to make. But if three different single nucleotide polymorphisms become one identity instead of three, we absolutely need to describe all three of those differences and attach them to the one identity that we present. So even when we decide that different things fall under one identity, all the information is still going to be discoverable, connected, and accessible.
Initially, we will choose our identity semantics and define them, then as new information arrives, we will decide, “Is this a new identity, or does this add to an existing identity?” We want to utilize as much computing power as we can, but at all times, our ultimate source of truth for those decisions is going to be humans who know the subject area best. That’s how we are going to deliver a trustworthy service.
Jeff: You want to create a clear, consistent, and comprehensive rule set upfront. For people who haven't tried to define entities and aggregate information, it feels very arbitrary, but we find that when you're organizing information, you want to err on the side of being clear and consistent. You can find nuances of the science in specific cases, but implementing rules for everything causes things to get scattered, and makes things harder to find and collect. Instead, it’s better to create a rule that works for 99% of everything. In the chemistry space, we sometimes get correspondence from a scientist who says, “I see you described this thing this way in CAS REGISTRY, and you're wrong, I have data that says it’s better described a different way,” and our usual answer is that our content is organized around discoverability. And that’s the point, the curation needs to consistently lead you to the data you’re looking for, even if it misses some scientific nuances along the way.
CAS: Are you using structure or function to create these identities for life sciences?
Mark: For many biological entities, especially proteins, it’s more function than structure, but it’s often a combination of both. In chemistry, we can easily create authority constructs based entirely on structure, but that’s not the same for life sciences. We have several different authorities we have to solve for—proteins, antibodies, organisms, etc. Each of them will need to be approached in their own unique way. We have to move away from just structure-based identities as we move into life sciences in a way that we have not had to in the past.
CAS: How do you plan to tackle the longevity of defining by function when biological function changes over time and with new research?
Jeff: We always try to future-proof things, and I don’t think we can predict where life sciences will be in ten years, but, as we create consistent constructs, we’ll build some flexibility and extensibility into that. The key to this will be recognizing when to keep using the current construct and when a new branch of science emerges that requires a new construct. People won’t stop wanting to describe proteins, but there might be a subcategory of some kind that arises that needs new authority constructs to be created.
Mark: I think that if you do a solid job of getting reasonable identity semantics at the beginning, then you set yourself up for a future where it's most likely that scientists will subdivide one of your identities into a few more specific identities. If you think about genus-species naming of organisms, it worked for a really long time. The need for sub-species didn't completely invalidate the genus-species model. I think we can set constructs up where they might get more specific about identity than when we started out, but they won’t become completely outdated.
CAS: What makes CAS the right organization to aggregate these fragmented life science authority constructs?
Jeff: There is some self-sustaining nature once you become a worldwide authority on something. If you look at chemical substances, nobody second-guesses CAS REGISTRY. It's the authoritative place. There are some other substance collections, but we are positioned as an organization in a way that most other organizations are not. We are part of the American Chemical Society, our mission isn't just to be profitable, our mission is to facilitate science. We have the people, the expertise, and the space to tackle this life science data and organize it in the way we did with CAS REGISTRY. I think anything short of us being the authoritative space for proteins when this project is complete is us falling short. That's what we're going to do.
Mark: The desire for agreement on a meaningful set of identities within life sciences is pretty universal, and it’s a problem we’ve already solved in small molecule chemistry. Life scientists are looking for a clear set of identities and the relationships between them to be defined so they can organize data around that structure. We see ourselves wading into that complexity and bringing order to the chaos. When we get far enough with that, people will accept that we know all the proteins that matter to most people and talk about them in a way that most people can make use of. When they don’t see their protein in our collection, they’ll bring it to our attention, so we can add it rather than starting their own entirely separate database. That’s what we’re aiming for—not just having an authority construct collection but being an authority within the life sciences space.
CAS: How is what you’re doing different than what’s already been done?
Mark: I think, in a lot of scientific spaces, people solve a problem for themselves and their colleagues, locally, in a way that works for them. They create a database that has the identities of things they care about, described in a way that they understand, and the information about those. Meanwhile, someone in a slightly overlapping area who's doing some of the same work, and some different work, is choosing different ways to talk about those identities and assembling different information. Then we start to see projects unite some of those smaller spaces to get all those people on the same page and facilitate discovery across those boundaries. So that's already happening in life sciences, but there hasn't been an approach to bring a large amount of it together with one uniting set of identities.
The goal is to get to a place where you can come in with your protein target, with no awareness that it is a member of three or four biological pathways, but when you've found that identity within this set of information, you now see all those connections. Additionally, you might also learn that it's being used as a biomarker for a disease state or treatment outcome. Being able to aggregate all the information from different places around one identity, which you can then find and use, creates value that wasn’t previously accessible.
Jeff:目前有许多小型组织在各自整合拼图的一小部分。例如,疾病本体论(Disease Ontology)的存在、Uniprot 对蛋白质的整理,以及 NCBI 的生物数据库。这些机构本身都是权威构建者,但它们之间缺乏有效的互联互通,用户必须在不同平台间来回切换。我们真正追求的是一个统一的数据集,让用户能够在此一站式获取并连接、协调整个生命科学领域的数据。为此,我们必须整合 Uniprot、Ensemble 和 CAS REGISTRY 对蛋白质的不同视角,将其协调为单一的视图和统一的标识体系。在此基础上,用户仍可回溯至所有原始的小型数据库。一旦实现互联,其价值将远超各部分之和,因为它能揭示以往无法发现的洞见。
Nicole:我们致力于成为全球最全面的生命科学数据库。我们希望解决客户面临的各种问题与挑战,而他们目前最迫切的需求正是生命科学领域的权威构建。
CAS:在进入这一领域时,你们是否遇到过来自现有数据库的阻力?
Mark:你不可能让所有人都满意。在某些节点上,你必须行使权威并做出抉择。要挑战那些根深蒂固的既有标准确实很难,我们也会尽量避免这样做。我们无意强迫所有人改变习惯的术语,但我们希望在现有基础上进行构建。
原则很简单:我们将与现有数据库保持一致。但在实践中,做出这些选择无疑更具挑战性。如果两个数据库对同一事物有不同的称呼,我们要么二选一,要么创造一个新术语并尝试优化两者。我知道这听起来很难,但我们认为,攻克这些难题正是我们为该领域创造的价值所在,也是我们希望为客户提供的核心成果。
Nicole:我们绝不想疏远那些使用现有数据库并从中获益的用户。我们希望做的是在现有基础上增加价值。我们尚未与相关策展人探讨过他们的看法,因此他们的态度还有待观察!
Jeff:实际上,从事数据策展的科学家只是少数,从某种程度上讲,他们并非我们最需要担心的群体,因为我们的目标是服务大多数人。我们从许多科学家那里了解到,他们获取所需数据时正面临重重困难。
Nicole:确实如此!最近我与一位创建了抗体信息数据库的科学家交谈。他说他这样做并非出于对整理抗体数据的兴趣,而是因为他需要这些数据来运行预测模型。我认为这代表了许多科学家的处境,数据获取障碍阻碍了研究进程,而这正是我们计划解决的问题。
CAS:如果能挥动魔杖解决权威构建中的一个问题,那会是什么?它将产生什么影响?
Mark:我希望解决的问题是让人们用相同的术语表达相同的含义。如果能做到这一点,一切都会简单得多。这项工作很大一部分在于梳理科学家描述事物所用的词汇,找到正确的标识,并将词汇与该标识关联起来。如果我们能统一语言并达成共识,就能省去这一繁琐过程。
Jeff:对我而言,我希望更多人能理解权威构建的意义,并具备相应的愿景与热情,以实用的方式去创造它们。即使在支持这一工作的组织内部,我仍需花费大量时间向他人解释其必要性与价值。虽然倡导这项工作很有意义,但这在一定程度上分散了我对核心工作的精力,而我最热衷的其实是处理数据和构建系统。






