English summary for screening — check the original posting before applying.
This role involves building and managing an in-house data utilization platform from scratch, covering everything from construction to user deployment and lifecycle operations. You will be part of a cross-functional team of data scientists, engineers, and project leaders.
Must-haves
- Experience with data modeling (dimensional modeling, data mart design)
- Experience developing and operating data infrastructure and ETL/ELT processes using Databricks
- Experience building data pipelines using Python (PySpark) for batch/streaming processing
- Experience establishing CI/CD deployment flows
- Experience with data governance and security implementation (access control, data quality checks)
- Experience supporting data users (data scientists, analysts, business departments)
Nice-to-haves
- Experience with unstructured data infrastructure expansion
- Experience with metadata and catalog management (table/column descriptions, lineage)
- Experience developing visualization applications (e.g., Streamlit)
- Experience building monitoring systems for data usage (query logs, lineage)
- Experience designing and operating mechanisms for promoting ad-hoc datasets to managed layers
- Experience detecting and managing unused tables/columns
- Experience designing automated data deletion using TTL
Tech stack
PythonSQLAWSGoogle Cloud PlatformDatabricksBigQueryTROCCOAmazon Aurora (MySQL)DockerECSTerraformGitHub ActionsPyTorchscikit-learn