Prof. Leser forscht an der automatisierten Analyse und Integration großer biomedizinischer Datenmengen, insbesondere zur Unterstützung personalisierter Krebstherapien. Seine aktuelle Arbeit konzentriert sich auf die Extraktion von Wissen aus medizinischer Literatur und klinischen Daten mittels Machine Learning und Large Language Models, um Ärzte bei der Diagnose und Therapieplanung zu unterstützen – etwa durch automatische Identifikation von Patienten mit ähnlichen molekularen Profilen oder durch intelligente Literaturkuration für Wissensdatenbanken. Daneben entwickelt er Methoden zur energieeffizienten Ausführung wissenschaftlicher Workflows auf Rechenclustern. Seine Lösungen adressieren Krankenhäuser, Biobanken und Pharmaunternehmen, die große Mengen an Genomik-, Pathologie- und Behandlungsdaten verwalten und analysieren müssen.
🔒 Das System hat 1271 mögliche Industrie-Partner gefunden — Firmen, Scores und Begründungen sind nur für eingeloggte Nutzer:innen sichtbar. Anmelden
Prof. Dr. Ulf Leser
HU-FIS-Profil ↗Helmholtz-Einstein International Berlin Research School in Data Science (HEIBRiDS)
other
GRK 2424: Computermethoden für personalisierte Therapien in der Onkologie
other
SFB 1404/2: FONDA – Grundlagen von Workflows für die Analyse großer naturwissenschaftlicher Daten
other
Förderer: Land Berlin - Andere Zeitraum: 07/2003 - 12/2007 Projektleitung: Prof. Dr. Ulf Leser
Zeitraum: 01/2007 - 12/2007 Projektleitung: Prof. Dr. Ulf Leser
Förderer: DFG Sonderforschungsbereich Zeitraum: 01/2009 - 12/2012 Projektleitung: Prof. Dr. Ulf Leser
Bioinformatics · DOI
MOTIVATION: Text mining has become an important tool for biomedical research. The most fundamental text-mining task is the recognition of biomedical named entities (NER), such as genes, chemicals and diseases. Current NER methods rely on pre-defined features which try to capture the specific surface properties of entity types, properties of the typical local context, background knowledge, and linguistic information. State-of-the-art tools are entity-specific, as dictionaries and empirically optimal feature sets differ between entity types, which makes their development costly. Furthermore, features are often optimized for a specific gold standard corpus, which makes extrapolation of quality measures difficult. RESULTS: We show that a completely generic method based on deep learning and statistical word embeddings [called long short-term memory network-conditional random field (LSTM-CRF)] outperforms state-of-the-art entity-specific NER tools, and often by a large margin. To this end, we compared the performance of LSTM-CRF on 33 data sets covering five different entity classes with that of best-of-class NER tools and an entity-agnostic CRF implementation. On average, F1-score of LSTM-CRF is 5% above that of the baselines, mostly due to a sharp increase in recall. AVAILABILITY AND IMPLEMENTATION: The source code for LSTM-CRF is available at https://github.com/glample/tagger and the links to the corpora are available at https://corposaurus.github.io/corpora/ . CONTACT: habibima@informatik.hu-berlin.de.
Lecture notes in computer science · DOI
The VLDB Journal · DOI