TG-Protein CLIP: Ontology-Guided Contrastive Learning for Protein Function Prediction via Sequence–Structure Fusion

Authors

  • Matěj Šimánek Faculty of Informatics, Czech University of Life Sciences Prague, Prague, 165 21, Czech Republic
  • Janek Novotný Faculty of Informatics, Czech University of Life Sciences Prague, Prague, 165 21, Czech Republic

DOI:

https://doi.org/10.64972/dea.2023.v2i1.3495d:57-70

Keywords:

Protein Function Prediction, Geometric Deep Learning, Ontology Prompting, Zero-Shot Recognition, Representation Learning

Abstract

Protein function prediction is still difficult for sequences with low similarity to known proteins, those lacking structural information, and those with sparse functional annotations in long-tail ontologies. Text-Guided Protein CLIP (TG-ProteinCLIP) is introduced in this paper as a new multi-modal framework that aligns fused sequence-structure representations with natural-language descriptions of molecular functions. Residue tokens from a protein language model are combined with geometric graph features via reliability-gated cross-attention, and ontology-aware prompts encode names, definitions, synonyms and hierarchical context. A confidence-calibrated contrastive objective supports both closed-set classification and zero-shot retrieval. According to experiments with temporally separated benchmark splits of 42,618 training proteins, 5,327 validation proteins and 6,104 test proteins, the proposed model achieved a protein-centric Fmax of 0.612 and an area under the precision-recall curve of 0.579. These values are 6.8% and 7.4% higher than the best sequence-only baseline. Remote-homology proteins with less than 20% sequence identity show an increase in Fmax from 0.421 to 0.503 and a decrease in expected calibration error from 0.087 to 0.041. Ablation and robustness experiments indicate that textual hierarchy, local geometric encoding and reliability gating offer complementary support for the gain. Therefore, it can be seen that language-guided alignment offers a practical way to achieve accurate, interpretable and label-efficient prediction of protein function under real-world annotation and structural uncertainty.

Downloads

Published

2023-02-19

How to Cite

Šimánek, M., & Novotný, J. (2023). TG-Protein CLIP: Ontology-Guided Contrastive Learning for Protein Function Prediction via Sequence–Structure Fusion. Data Engineering and Applications, 2(1), 5d:57–70. https://doi.org/10.64972/dea.2023.v2i1.3495d:57-70

Issue

Section

Articles

Similar Articles

<< < 2 3 4 5 6 7 8 > >> 

You may also start an advanced similarity search for this article.