Data Engineer (Working Student)
Allianz Commercial · Munich, Germany
Engineered a Streamlit reporting application integrated with SQL and Azure Databricks pipelines, cutting report delivery time by 40%. Automated Python- and Docker-based ETL workflows for scheduled data processing, eliminating ~8 hours of manual effort per week. Built scalable PySpark data transformation pipelines in Azure Databricks for reporting and analytics use cases. Designed and deployed a production LLM-based email-processing pipeline using Claude Opus and the Anthropic API to extract and normalize unstructured email content into structured JSON. Orchestrated data processing workflows using YAML-based pipeline configuration and Azure Databricks/PySpark. Applied validation and data quality checks to ensure reliable pipeline outputs and accurate downstream reporting. Used AI-assisted development tools, including Databricks AI Assistant, to accelerate PySpark development, debugging and solution iteration.