Anna Shadrova forscht an korpuslinguistischen Methoden und Infrastrukturen für die Analyse von Sprachkontakt und grammatischen Veränderungen. Sie entwickelt und pflegt Korpusressourcen, modelliert Analyseebenen formal und setzt Verfahren des maschinellen Lernens ein, um sprachliche Variation und Dynamiken in mehrsprachigen Kontexten zu untersuchen. Ihre Arbeit liefert Unternehmen und Institutionen, die mit mehrsprachigen Daten arbeiten — etwa in Sprachverarbeitung, Lokalisierung oder linguistischer Analyse — standardisierte Dateninfrastrukturen und automatisierte Analyseverfahren. Die Forschung ist relevant für Sprachtechnologie, Sprachdokumentation und vergleichende Linguistik.
🔒 Das System hat 190 mögliche Industrie-Partner gefunden — Firmen, Scores und Begründungen sind nur für eingeloggte Nutzer:innen sichtbar. Anmelden
Anna Shadrova
HU-FIS-Profil ↗Das Teilprojekt Korpuslinguistische Methoden der Forschergruppe Emerging Grammars (RUEG 2) beschäftigt sich zum einen mit der Pflege und Weiterentwicklung der Korpusressourcen und der Infrastruktur innerhalb der Forschergruppe. Zum zweiten erarbeitet und unterstützt es die konzeptuelle und formale Modellierung der unterschiedlichen Analyseebenen. Des weiteren setzt es Verfahren des maschinellen Lernens ein, um die Daten auszuwerten.
Programme for international student assessment/Internationale Schulleistungsstudie · DOI
This report provides a systematic review and empirical evidence related to the experiences of middle-income countries and economies participating in the Programme for International Student Assessment (PISA), 2000 to 2015. PISA is a triennial survey that aims to evaluate education systems worldwide by testing the skills and knowledge of 15-year-old students. To date, students representing more than 70 countries and economies have participated in the assessment, including 44 middle-income countries, many of which are developing countries receiving foreign aid. This report provides answers to six important questions about these middle-income countries and their experiences of participating in PISA: What is the extent of developing country participation in PISA and other international learning assessments? Why do these countries join PISA? What are the financial, technical, and cultural challenges for their participation in PISA? What impact has participation had on their national assessment capacity? How have PISA results influenced their national policy discussions? And what does PISA data tell us about education in these countries and the policies and practices that influence student performance? The findings of this report are being used by the OECD to support its efforts to make PISA more relevant to a wider range of countries, and by the World Bank as part of its on-going dialogue with its client countries regarding participation in international large-scale assessments.
Journal of Data Mining & Digital Humanities · DOI
The social sciences and digital humanities have recently adopted the machine learning technique of topic modeling to address research questions in their fields. This is problematic in a number of ways, some of which have not received much attention in the debate yet. This paper adds epistemological concerns centering around the interface between topic modeling and linguistic concepts and the argumentative embedding of evidence obtained through topic modeling. It concludes that topic modeling in its present state of methodological integration does not meet the requirements of an independent research method. It operates from relevantly unrealistic assumptions, is non-deterministic, cannot effectively be validated against a reasonable number of competing models, does not lock into a well-defined linguistic interface, and does not scholarly model topics in the sense of themes or content. These features are intrinsic and make the interpretation of its results prone to apophenia (the human tendency to perceive random sets of elements as meaningful patterns) and confirmation bias (the human tendency to perceptually prefer patterns that are in alignment with pre-existing biases). While partial validation of the statistical model is possible, a conceptual validation would require an extended triangulation with other methods and human ratings, and clarification of whether statistical distinctivity of lexical co-occurrence correlates with conceputal topics in any reliable way.
Die korpuslinguistische Arbeit untersucht den Erwerb von Koselektionsbeschränkungen bei Lerner*innen des Deutschen als Fremdsprache in einem quasi-longitudinalen Forschungsdesign anhand des Kobalt-Korpus. Neben einigen statistischen Analysen wird vordergründig eine graphbasierte Analyse entwickelt, die auf der Graphmetrik Louvain-Modularität aufbaut. Diese wird für diverse Subkorpora nach verschiedenen Kriterien berechnet und mit Hilfe verschiedener Samplingtechniken umfassend intern validiert. Im Ergebnis zeigen sich eine Abhängigkeit der gemessenen Modularitätswerte vom Sprachstand der Teilnehmer*innen, eine höhere Modularität bei Muttersprachler*innen, niedrigere Modularitätswerte bei weißrussischen vs. chinesischen Lerner*innen sowie ein U-Kurven-förmiger Erwerbsverlauf bei weißrussischen, nicht aber chinesischen Lerner*innen. Unterschiede zwischen den Gruppen werden aus typologischer, kognitiver, diskursiv-kultureller und Registerperspektive diskutiert. Abschließend werden Vorschläge für den Einsatz von graphbasierten Modellierungen in kernlinguistischen Fragestellungen entwickelt. Zusätzlich werden theoretische Lücken in der gebrauchsbasierten Beschreibung von Koselektionsphänomenen (Phraseologie, Idiomatizität, Kollokation) aufgezeigt und ein multidimensionales funktionales Modell als Alternative vorgeschlagen.