Analyzing academic career trajectories through Process Mining (PM) methods offers critical insights into organizational behavior and promotion patterns, yet such studies are frequently hindered by the labor-intensive nature of manual data collection. Academic staff information is typically fragmented across institutional directories, personal webpages, and curricula vitae (CVs) that vary significantly in structure and accessibility. This study presents a semi-automated Python-based framework designed to identify, retrieve, and transform this unstructured information into structured event logs ready for process analysis. The system implements a modular pipeline that combines web scraping with Large Language Models (LLMs) to interpret heterogeneous textual sources. The LLM component serves as a core enrichment engine, parsing complex CVs and personal profiles to extract essential career milestones, such as educational history, appointments, and publication records, into a normalized JSON schema. This structured data is subsequently converted into CSV format, enabling integration with process mining platforms such as Celonis to identify advancement bottlenecks and career progression stages. By bridging the gap between scattered public records and systematic analytical tools, this infrastructure creates a practical method for large-scale academic career modeling. The proposed approach significantly reduces manual effort while maintaining data integrity through a deterministic validation layer. Across five Israeli academic institutions, the pipeline processed over 9,000 faculty profiles and, following human validation of 390 candidate records, yielded 106 validated new cases. These records expanded the existing dataset from 373 to 479 cases across four institutions, demonstrating the scalability and practical value of the proposed approach for process-oriented research in higher education management.
Keywords
Process Mining, Large Language Models, Data Extraction, Academic Career Trajectories, Web Scraping.