Showing posts with label semantics. Show all posts
Showing posts with label semantics. Show all posts

2024-06-27

The Variabilities of Dopamine (₯) - PART I:ChEBI:18243

 


Fig 1: dopamine chemical families
Upon reading the article, the Variabilities of Dopamine (₯) - Prequel, we discover that our understanding of dopamine encompasses various facets, which shift with the passage of time and the context in which we perceive them. Dopamine is usually introduced by explaining that it is an organic chemical from the catechol/catecholamine and phenethylamine families. But wait, what are the names of all these chemical compounds popping up all of a sudden? What's going on with this line of scientific family tree? For complex information, we can use a powerful tool that integrates the breadth and depth of knowledge units, the so-called Knowledge Graph (KG) and an ontology that clarifies the structure and "relationships" between knowledge. The ontology mentioned here was originally established by scholars from various disciplines. It provides a method for computers to "understand" the semantics of data to perform calculations, describe it in a universal machine-readable format, and integrate other knowledge sources for application in different fields.


The most familiar KG is the information box displayed on the right side of Google search results. As a framework for understanding the entities and their relationships in the real world, the information that appears most often in the structured database - Wikipedia, and one of the wisdom behind this comes from semantic & ontological structures. If the ontology is placed in the context of popular science, it will help us quickly grasp the information, which is the saying that a picture is worth a thousand words. Because ontology has the advantage of visualizing data and combining images and textual knowledge trees to help readers move from familiar visual representations to abstract scientific understanding, we will try to use different ontologies to convey the complex stories of dopamine and bring readers to the concept of dopamine from the micro and macro perspectives respectively, and understand "her" properties, structure and basic knowledge more accurately.

Ontology: ChEBI

Within the realm of biomedical ontology, every knowledge unit—whether an entity or a node—embodies a distinct biological or chemical entity (such as genes, cells, diseases, drugs, and more), revealing essential attributes inherent to each compound. In the ontology, each connected edge (relation) describes the interaction or relationship between entities (such as "drugs treat diseases" or "cells up-regulate genes"). Therefore, using the standard vocabulary of "Chemical Entities of Biological Interest (ChEBI)" as a formal introduction to dopamine is quite consistent with our understanding of dopamine's molecular structure, role in biology, and common understanding among entities. Chemistry-specific relationships (such as the relationship types and family pedigrees depicted in Figure 2) take into account the need for both broad and in-depth understanding.

ChEBI is part of the Open Biomedical Ontology (OBO) of the Heidelberg-based European Molecular Biology Laboratory- European Bioinformatics Institute (EMBL-EBI) and focuses on integrating and describing data on "small" compounds. "Molecular entity" refers to any structurally or isotopically distinct atom, molecule, ion, ion pair, free radical, radical ion, complex, conformational isomer, etc. Recognizable as individually distinguishable entities, the molecular entities in question are either natural products or synthetic products designed to intervene in biological processes. ChEBI contains the relationship between an ontology classification and a specified molecular entity, or an entity class and its parents and/or children, usually called as a "parent-child relationship."

 Figure 2: part knowledge structure of dopamine

Dopamine in ChEBI Ontology
Dopamine (ChEBI: 18243) is defined in ChEBI as: Catechol in which the hydrogen at position 4 is substituted by a 2-aminoethyl group. 18243 is the "unique identification number" provided by ChEBI to each entity. In Figure 2, ChEBI: 18243 is a family of main group molecular entities (ChEBI: 33579), including one or more of any group in groups 1, 2, 13, 14, 15, 16, 17 and 18 of the periodic table of elements. A molecular entity of atoms. Dopamine, in this tree view, can be described as being a (is a) catecholamine (ChEBI: 33567), a monoamine molecular messenger (ChEBI: 25375), and an organic molecular entity (ChEBI: 50860).


However, dopamine also belongs to another branch of the family of chemical entities: phenols (ChEBI: 33853). As examined in Figure 3, it can be seen that dopamine is also a polyatomic entity (ChEBI: 36357), an organic aromatic compound (ChEBI: 33659), and a catechol (ChEBI: 33566). In Figure 1, it can also be seen that the conjugate base and conjugate acid of dopamine is dopamine (1+) (ChEBI: 59905), which is any mammalian "metabolite" produced during human metabolic reactions. In biology, the scientific role is related to neurotransmitter disorders. In subsequent articles, we will use other ontologies to further introduce it.

Figure 3: part knowledge structure of dopamine (graph view)

In one of the main relationship categories of ChEBI "has role", the biological role of dopamine is described (ChEBI: 24432), including β-adrenergic agonists, E. coli metabolites, dopaminergic drugs, mimetic sympathogenic agents, mouse metabolites, and human metabolites. The application relationship (ChEBI: 33232) describes the intended use of the molecular entity or its parts by humans. Therefore, the application level of dopamine includes: cardiotonic drugs, β-adrenergic agonists, dopaminergic drugs and sympathomimetic agents, etc. 

 

Figure 4: similar structures & has part relations within dopamine
There are also descriptions and links in the ontology. If we use the "has part" relationship query in ChEBI, we can also find compounds containing dopamine structures, which currently include at least 1387 entities, and we can also find 59 compounds similar to this structure, as shown in Figure 4 shown. 

The latest version (2024/06/27) of ChEBI contains nearly 62,000 compounds and more than 190,000 "relationships". ChEBI has a wide range of applications, including the construction of biomedical knowledge graphs, its rich hierarchical structure and other relationship types, which can provide therapeutic assistance in identifying chemical entities in Alzheimer's disease and dementia literature. So the final question is, is it feasible to directly use ChEBI’s dopamine to communicate directly with the public? 

Of course! As shown in Figure 5, PubChem, an open chemical database of the U.S. National Library of Medicine and the National Institutes of Health (NIH), uses ChEBI: 18243 as the source of information describing dopamine. In addition, when the news media acquaints us with dopamine’s role in pain, they cite ChEBI: 18243 as their source. ChEBI’s portrayal of dopamine reveals it as a chemical entity. But what narrative lies within its ‘cell family lineage’? (To be continued…)"


Fig 5: Using ChEBI:18243 for the reference source in PubChem


2020-06-11

A brief introduction of Dryad in 10 minutes

Update information (2020-10-15): 

  •  the Data package DOI + Version numbers are retiring 
  •  the Data Paper in PDF file & OAI-PMH (DSpace) are retiring, replaced by API (https://datadryad.org/api/v2/docs/) ,
  • the schema.org embedded metadata ( HTML landing page /JSON-LD), and DataCite metadata export. 

Acknowledgements:

We thank Daniella Lowenberg, Dryad Product Manager, University of California Curation Center (UC3) for useful clarifications and comments.

Andrea Wei-Ching Huang. (2020). The Story of Research Data Repository (RDR) (Version 2020-10-15T08:08:30.049287) [Data set]. Retrieved from https://data.depositar.io/dataset/rdr-story
 
resource: https://data.depositar.io/dataset/rdr-story/resource/266a3d4e-217d-408f-ada3-14e540efaea5

 

2018-10-27

Revisit: Reuse of Structured Data: Semantics, Linkage, and Realization (2)




(continue from part I) / Library and Information Science, 43.1 (2017): 7-46. / [[中文]]


RESEARCH HIGHLIGHTS: 

# An old record is not a data but now defined as a new semantic dataset. 
i.e. its triples, graphs, links, file formats ...
i.e. its revised, vocabulary encoded versions ...
ex. data:d2148340 a dcat:dataset.
#files:json-ld, ttl, XML

#
A new method to curate, publish & visualize LOD graphs via CKAN portal.
i.e. two models for one dataset published in two views.
ex. data:d2148340 a dcat:dataset. # Dublin Core @schema1
ex. data:d2148340 a data:Refined. # more semantics@schema2

#
Validation & Reproducibility: Provenance and Contexts are in details.

Practices

Example: data:d2148340 (click to enlarge)
We then make use of structured records (XML files) from a digital archive catalogue, and convert the records into semantically rich and interlinked resources on the Web. This is realized as a unified Linked Data catalogue to several digital archive collections. Our work results in a LOD catalogue (data.odw.tw) available to the public at the website . The following five parts are involved in realizing this website. 


A catalogue record, about a species of Pleione Formosana (data:d2148340), is used throughout in the paper as an example to demonstrate the way we model, convert, and represent the semantics of a structured record.

R4R Ontology (click to enlarge)
Part 1: Exploring data reuse relations in a shared context -- We review our previous research about the Relation for Reuse Ontology (R4R). In particular, we provide mechanisms for reusing article, data, and code with some flexibility of encoding provenance and license information.

Part 2: Comparing two different data conversion approaches to providing LOD for an archive catalogue -- We show two different scenarios: (1) The LOD catalogue is converted directly from a relational database, and (2) the LOD catalogue is generated from a series of format conversions --- from XML to CSV, and then to RDF. 

KB links Example (click to enlarge)
Part 3: Data profiling, cleaning and mapping -- We demonstrate format conversion processes, and we discuss the pros and cons of various ways in handling broken links in source datasets. In addition, we mapped and linked catalogue records to three external knowledge bases: GeoNames, Wikidata, and Encyclopedia of Life.  

Part 4: Using CKAN as a Linked Data platform -- We briefly introduce CKAN, an open source web-based data portal software package for curating and publishing datasets. CKAN provides data preview, search, and discovery, especially with regard to geospatial datasets. We built several extensions to CKAN in order to deposit, publish, browse, and search Linked Data. Various Linked Data representations of a catalogue record --- Turtle, RDF/XML, and JSON-LD --- can all be downloaded and reused.

Part 5: Designing an ontology for data representation and reuse -- We design an ontology voc4odw which includes the following 3 modules:

(1) The Core Model. It is comprise of a data model and a conceptual model. 




The data model represents key data structure and relation. It is a framework to illustrate data source,derivation, and provenance.

The voc4odw Data Model (click)
The conceptual model incorporates Simple Knowledge Organization System (SKOS); it also connects to key event concepts. The conceptual model allows for data contextualization using common and domain knowledge vocabularies.



(2) The Curation Model. It is responsible for disclosing the identification, classification, and publication of structured records at a curation platform, such as the classification of themes, the assignment of data identifiers, and the publication of datasets.

(3) A vocabulary voaf:Vocabulary. It is defined as "A vocabulary used in the Linked Data cloud", from the Vocabulary of a Friend . This module is to relate the Core Model to external common vocabularies. Some hierarchy relations between different external vocabularies can be traced with this vocabulary.


voc4odw ontology
Common Knowledge
Prefix
Namespace
Description
cc
http://creativecommons.org/ns#
csvw
http://www.w3.org/ns/csvw#           
dc
dcat
dct
5.       DCMI Metadata Terms
dctype
http://purl.org/dc/dcmitype/
6.       DCMI Type Vocabulary
event
http://purl.org/NET/c4dm/event.owl#
7.       Event Ontology
foaf
geo
http://www.w3.org/2003/01/geo/wgs84_pos#
gn
10.     GeoNames Ontology
gns
11.     GeoNames Entity
lcsh
http://id.loc.gov/authorities/subjects
org
prov
r4r
schema
16.     Schema.org
skos
time
http://www.w3.org/2006/time#
18.     W3C  Time Ontology
voaf
http://purl.org/vocommons/voaf#
wde
http://www.wikidata.org/entity/
20.     Wikidata Entity
 Domain Knowledge
aat
http://vocab.getty.edu/aat/
dwc
2.       Darwin Core Terms
dwciri
3.       Darwin Core terms
eol
4.       The Encyclopaedia of Life (EOL)
txn
http://lod.taxonconcept.org/ontology/txn.owl#
Local Namespace
voc
http://voc.odw.tw/ontology#  
agent
article
code
data
5.      Linked Data for ODWeb
evt84
6.      Event Entity in ODW
project
7.      Project Entity in ODW
r1 (n)
http://data.odw.tw/r1/   (r2, r3…)
refined
http://data.odw.tw/refined/
catdat
http://catalog.digitalarchives.tw/

2016-10-04

LOD can be hugged by human: data.odw.tw.





LODs are nowadays all in love with machines (try 614 examples represented with 3 event types: schema:CreateAction, schema:OrganizeAction, and schema:PublicationEvent.)  Yet it can be hugged by human through CKAN at data.odw.tw
This is the Prat II of the Story of One Leaf, an implementation of the R4R ontology for reusing digital objects for 843312 reused objectsFor instance, see the leaf Pleione formosana Hayata (台灣一葉蘭) in the reusing context.


The Story of One Leaf/ D Version Examples: data:d2148340 (Pleione formosana Hayata 台灣一葉蘭)
data:d2148340 a data:Reused, r4r:RRObject, dcat:Dataset ;
r4r:hasProvenance data:p20160530-d2148340 ;
dc:publisher "中央研究院生物多樣性研究中心"^^rdf:PlainLiteral ;
dc:source "台灣本土植物資料庫
(http://taiwanflora.sinica.edu.tw/)"^^rdf:PlainLiteral ;
dc:date "採集日期:1993-04-25"^^rdf:PlainLiteral ;
dc:coverage "國家:台灣"^^rdf:PlainLiteral,
"最低海拔:1650"^^rdf:PlainLiteral,
"行政區:宜蘭縣大同鄉"^^rdf:PlainLiteral .
  • The case data:d2148340 in details is a data:Reused resource, a basic data component in the current Linked Data in the Open Data Web. It uses the R4R ontology to declare its URI as http://data.odw.tw/record/d2148340, and thus it can be shared later with its refined semantic versions using the same URI (ex. its R1 Version). It has provenance information using r4r:hasProvenance to relate the r4r:RRObject with its provenance information.
  • The data:Reused resource is a dcat:Dataset which is a collection of triples, published and curated by ODW, and available for access or download in many versions and formats.
  • Currently, the data:Reused resources are described with Dublin Core 15 Elements (DC15) in close relation to  their primary data. Here we call them DC 15 Version or D Version. Further semantic refinement is done by the refined versions, defined as data:Refined (we call them R Version, with different R1, R2, R3...) extracting information from the values of DC 15 properties. 



    See the leaf  in a semantically refined context:






The Story of One Leaf/ R Version Examples: data:d2148340 (Pleione formosana Hayata 台灣一葉蘭)
data:d2148340 a data:Refined, r4r:Data, dcat:Dataset ;
r4r:hasProvenance data:p20160706-d2148340 ;
txn:hasEOLPage eol:1134120;
dct:requires evt84:event-d2148340, evt84:phyCre-d2148340 ;
dcat:landingPage r1:r1-r2148340;
dcat:themeTaxonomy data:Biology .
evt84:phyCre-d2148340 a schema:CreateAction ;
event:factor dct:PhysicalResource ;
event:product dwc:PreservedSpecimen ;
dwc:eventDate "1993-04-25" ;
skos:inScheme dwc:HumanObservation ;
skos:scopeNote "specimen collection process" .
  • The data:d2148340 is a data:Reused, but also a data:Refined in different contexts. It is a semantically refined version (a r4r:Data in the R Version) of the data:Reused (a r4r:RRObject in D Version). In other words, data:d2148340 and r1:r1-r2148340 share the same URI according to the definition of the r4r:RRObject. which conveys the relation that the r1:r1-r2148340 (a r4r:Data) is a subclass of the data:d2148340 (a r4r:Object).
  • What distinguishes D Version and R Version most ?
    (1) the former with literal values, the latter with URI resource values.
    (2) the latter heavily relies on the event model (further described by using the former that the subject  dct:requires events to denote its people-place-time relations (extracting values of dc:coverage and dc:date from the former). Ex. the event evt84:phyCre-d2148340 triples as shown in the example.
    (3) the latter based on the curation context (data:ConceptTheme, ex. data:Biology) to enrich the resource with domain knowledge like EOL pages, dwc:PreservedSpecimenl and dwc:HumanObservation.




The following slides are a general overview for data.odw.tw.




Triples of the Girl Lost in Thought for Human:

Version D: represented with DC 15 Elements and Provenance (PROV-O)
Version R(r1): represented with refined spatial temporal information and linked to Wikidata, GeoNames, AAT …)