Fabian Lehmann forscht aktuell an zwei parallelen Schwerpunkten: (1) Mathematische Grundlagen für computergestützte Existenzbeweise von speziellen geometrischen Strukturen (Calabi-Yau und G2-Mannigfaltigkeiten), und (2) Verwaltung und Optimierung von wissenschaftlichen Datenanalyse-Workflows in verteilten Rechensystemen. Im Workflow-Bereich entwickelt er Methoden zur effizienten Ressourcennutzung, automatischen Performance-Vorhersage von Rechentasks und standardisierten Beschreibungssprachen für komplexe Analysepipelines — ein Problem, das besonders für Erdbeobachtung, Bioinformatik und große Datenmengen relevant ist. Seine Arbeiten adressieren die praktische Herausforderung, dass wissenschaftliche Workflows heute oft mit statischen, ungenauen Ressourcenschätzungen arbeiten, was zu Verschwendung oder Verzögerungen führt.
🔒 Das System hat 211 mögliche Industrie-Partner gefunden — Firmen, Scores und Begründungen sind nur für eingeloggte Nutzer:innen sichtbar. Anmelden
Ph.D. Fabian Lehmann
HU-FIS-Profil ↗Förderer: DFG Eigene Stelle (Sachbeihilfe) Zeitraum: 03/2026 - 02/2028 Projektleitung: Ph.D. Fabian Lehmann
Continuous integration and deployment are established paradigms in modern software engineering. Both intend to ensure the quality of software products and to automate the testing and release process. Today's state of the art, however, focuses on functional tests or small microbenchmarks such as single method performance while the overall quality of service (QoS) is ignored. In this paper, we propose to add a dedicated benchmarking step into the testing and release process which can be used to ensure that QoS goals are met and that new system releases are at least as "good" as the previous ones. For this purpose, we present a research prototype which automatically deploys the system release, runs one or more benchmarks, collects and analyzes results, and decides whether the release fulfills predefined QoS goals. We evaluate our approach by replaying two years of Apache Cassandra's commit history.
l L a b o r a t o r y } }
arXiv (Cornell University) · DOI
Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enabling the publication of workflows and their associated products according to the FAIR principles. This document reports on discussions and findings from the 2022 international edition of the Workflows Community Summit that took place on November 29 and 30, 2022.