Prof. Klein entwickelt statistische Methoden für die Analyse komplexer medizinischer Daten, die über einfache Mittelwertbetrachtungen hinausgehen. Ihr aktueller Fokus liegt auf multivariaten Verteilungsregressionsmodellen (Boosting Copulas) und Bayesianischen Machine-Learning-Ansätzen, um gleichzeitig mehrere klinische Endpunkte und ihre Abhängigkeiten realistisch abzubilden. Dies ermöglicht es Unternehmen im Gesundheitswesen und der Pharmaforschung, aus großen, komplexen Patientendatensätzen präzisere Erkenntnisse zu gewinnen — etwa zur Identifikation von Risikofaktoren in Subpopulationen oder für einzelne Patienten. Die Methoden sind besonders für die digitale Medizin, klinische Studienauswertung und personalisierte Diagnostik relevant.
🔒 Das System hat 425 mögliche Industrie-Partner gefunden — Firmen, Scores und Begründungen sind nur für eingeloggte Nutzer:innen sichtbar. Anmelden
Prof. Dr. Nadja Klein
HU-FIS-Profil ↗Im Zeitalter der Digitalisierung liegen vielen wissenschaftlichen Studien immer größere und komplexere Datenmengen zugrunde. Diese „Big Data“-Anwendungen bieten viele Ansatzpunkte für die Weiterentwicklung von statistischen Methoden, die insbesondere genauere und an deren Komplexität angepasste Modelle sowie die Entwicklung verbesserter Inferenzmethoden erfordern, um potentiellen Modellfehlspezifikationen, verzerrten Schätzern und fehlerhaften Folgerungen und Prognosen entgegenzuwirken. Das hier vorgeschlagene Projekt wird statistische Methoden für flexible univariate und multivariate Regressionsmodelle und deren genaue und effiziente Schätzung entwickeln. Genauer sollen durch einen probabilistischen Ansatz zu klassischen Verfahren des maschinellen Lernens effizientere und statistische Lernalgorithmen zur Schätzung von Modellen mit großen Datensätzen erarbeitet werden. Um die Modellierung der gesamten bedingten Verteilung der Zielgrößen zu ermöglichen, sollen darüber hinaus neuartige Verteilungsregressionsmodelle entwickelt werden, welche sowohl die Analyse univariater als auch multivariater Zielgrößen erlauben und gleichzeitig interpretiere Ergebnisse liefern. In all diesen Modellen sollen außerdem die wichtigen Fragen der Regularisierung und Variablenselektion betrachtet werden, um deren Anwendbarkeit auf Problemstellungen mit einer großen Anzahl an potentiellen Prädiktoren zu gewährleisten. Auch die Entwicklung frei verfügbarer Software sowie Anwendungen in den Natur- und Sozialwissenschaften (wie zum Beispiel zu Marketing, Wettervorhersagen, chronischen Krankheiten und anderen) stellen einen wichtigen Bestandteil des Projekts dar und unterstreichen dessen Potential, entscheidend zu wichtigen Aspekten der modernen Statistik und Datenwissenschaft beizutragen.
PLoS ONE · DOI
Although regression models play a central role in the analysis of medical research projects, there still exist many misconceptions on various aspects of modeling leading to faulty analyses. Indeed, the rapidly developing statistical methodology and its recent advances in regression modeling do not seem to be adequately reflected in many medical publications. This problem of knowledge transfer from statistical research to application was identified by some medical journals, which have published series of statistical tutorials and (shorter) papers mainly addressing medical researchers. The aim of this review was to assess the current level of knowledge with regard to regression modeling contained in such statistical papers. We searched for target series by a request to international statistical experts. We identified 23 series including 57 topic-relevant articles. Within each article, two independent raters analyzed the content by investigating 44 predefined aspects on regression modeling. We assessed to what extent the aspects were explained and if examples, software advices, and recommendations for or against specific methods were given. Most series (21/23) included at least one article on multivariable regression. Logistic regression was the most frequently described regression type (19/23), followed by linear regression (18/23), Cox regression and survival models (12/23) and Poisson regression (3/23). Most general aspects on regression modeling, e.g. model assumptions, reporting and interpretation of regression results, were covered. We did not find many misconceptions or misleading recommendations, but we identified relevant gaps, in particular with respect to addressing nonlinear effects of continuous predictors, model specification and variable selection. Specific recommendations on software were rarely given. Statistical guidance should be developed for nonlinear effects, model specification and variable selection to better support medical researchers who perform or interpret regression analyses.
Journal of Statistical Software · DOI
Over the last decades, the challenges in applied regression and in predictive modeling have been changing considerably: (1) More flexible regression model specifications are needed as data sizes and available information are steadily increasing, consequently demanding for more powerful computing infrastructure. (2) Full probabilistic models by means of distributional regression - rather than predicting only some underlying individual quantities from the distributions such as means or expectations - is crucial in many applications. (3) Availability of Bayesian inference has gained in importance both as an appealing framework for regularizing or penalizing complex models and estimation therein as well as a natural alternative to classical frequentist inference. However, while there has been a lot of research on all three challenges and the development of corresponding software packages, a modular software implementation that allows to easily combine all three aspects has not yet been available for the general framework of distributional regression. To fill this gap, the R package bamlss is introduced for Bayesian additive models for location, scale, and shape (and beyond) - with the name reflecting the most important distributional quantities (among others) that can be modeled with the software. At the core of the package are algorithms for highly-efficient Bayesian estimation and inference that can be applied to generalized additive models or generalized additive models for location, scale, and shape, or more general distributional regression models. However, its building blocks are designed as
Statistical Methods in Medical Research · DOI
We present a new procedure for enhanced variable selection for component-wise gradient boosting. Statistical boosting is a computational approach that emerged from machine learning, which allows to fit regression models in the presence of high-dimensional data. Furthermore, the algorithm can lead to data-driven variable selection. In practice, however, the final models typically tend to include too many variables in some situations. This occurs particularly for low-dimensional data ([Formula: see text]), where we observe a slow overfitting behavior of boosting. As a result, more variables get included into the final model without altering the prediction accuracy. Many of these false positives are incorporated with a small coefficient and therefore have a small impact, but lead to a larger model. We try to overcome this issue by giving the algorithm the chance to deselect base-learners with minor importance. We analyze the impact of the new approach on variable selection and prediction performance in comparison to alternative methods including boosting with earlier stopping as well as twin boosting. We illustrate our approach with data of an ongoing cohort study for chronic kidney disease patients, where the most influential predictors for the health-related quality of life measure are selected in a distributional regression approach based on beta regression.