About the Talk
Presenter
Serbulent Unsal
Serbulent Unsal completed undergraduate studies in Statistics and Computer Science at Karadeniz Technical University, then continued with a master’s degree in Medical Informatics at Middle East Technical University. During the master’s program, the work focused on multiscale computational tumor modeling, developing a tumor progression model using cellular automata and partial differential equations, under the supervision of Dr. Aybar Can Acar. In 2014, Ph.D. studies began in the same department, developing deep learning models for low-data protein function prediction: a thesis that is also part of a larger research project on discovering novel immune escape mechanisms and drug repurposing against them. Now completing the Ph.D., Serbulent Unsal works as a Senior Machine Learning Engineer at Antiverse, designing antibodies using machine learning and deep learning models.
Abstract
Proteins are macromolecules essential for life. Understanding and manipulating biological mechanisms requires understanding protein function, which is possible by studying the relationship between amino acid sequences, 3D structures, and function. Until now, only a small percentage of proteins have been functionally characterized (currently around 0.5% according to UniProt) due to the cost and time requirements of wet-lab-based procedures. Recently, protein function prediction (PFP) (the annotation of proteins with functional descriptions using statistical/computational methods) has gained importance for discovering the uncharacterized protein space and protein variants carrying functional changes. Among the many algorithmic approaches proposed to date, machine learning (ML), and deep learning (DL) in particular, have become popular in PFP due to their high predictive performance. The input data used by these ML/DL methods are numerical feature vectors representing the protein (i.e., protein representations), most often generated from the amino acid sequences available in databases such as UniProt. This talk evaluates protein representation methods for predicting the functional properties of proteins, comparing them across 4 challenging tasks. In total, 23 protein representation methods were evaluated, including classical approaches and advanced representation-learning methods. Finally, the talk introduces PROBE (Protein RepresentatiOn BEnchmark), an open-access tool with which users can evaluate new protein representation models on the benchmark tasks.
Date: 6 July 2022 – 6:00 PM (GMT+3)
Language: English