GCP Data Engineer - VOIS
Who we are
About this Role
What you will do
- Design, develop and maintain ETL and ELT pipelines that ingest, transform and load large datasets from multiple sources into GCP data platforms.
- Develop data solutions using BigQuery, Cloud Storage, Cloud Composer, Dataflow, Pub/Sub and Cloud Functions.
- Optimise data pipelines to improve performance, reliability and scalability.
- Develop and manage data models, schemas and storage solutions aligned with GCP good practices.
- Build effective and scalable Python code and data-processing scripts.
- Automate data workflows and scheduled processes using appropriate orchestration tools.Implement data validation, cleansing and quality controls to maintain data accuracy and integrity.
- Monitor data pipelines, identify issues and support the timely delivery of data.Optimise BigQuery performance and manage storage and processing costs.
- Develop and maintain Kafka streaming integrations using the Python SDK.Build and integrate back-end components that improve application responsiveness and performance.
- Collaborate with data scientists, analysts and business stakeholders to define and deliver data requirements.
- Coordinate with internal teams and external partners to support effective project delivery.
- Apply a detail-focused and solution-oriented approach while considering wider business objectives.
Who you are
- You have strong practical experience with Python, including the development and deployment of data-processing scripts.
- You have hands-on expertise in Google Cloud Platform, particularly BigQuery, Cloud Storage, Cloud Functions, Cloud Composer and Pub/Sub.
- You have experience designing, developing and maintaining ETL or ELT data pipelines.You are proficient in SQL for querying, transforming and processing data.
- You understand data modelling, schemas, data storage and data quality principles.
- You are familiar with workflow orchestration tools such as Apache Airflow.
- You have knowledge of version control systems such as Git.
- You have experience with, or knowledge of, Kafka-based streaming solutions and Python SDK integrations.
- Familiarity with Teradata and PowerDesigner would be beneficial.
- You communicate clearly in writing and in conversation, including when working with clients and business stakeholders.
- You collaborate effectively across technical and non-technical teams.
- You can understand current and future business objectives and translate them into effective technology solutions.
- You demonstrate sound judgement, ownership and a collaborative approach to project delivery.
Not a Perfect Fit?
What’s in it for you
- The opportunity to work with modern GCP data services and cloud-native data architectures.
- Exposure to large-scale data engineering projects with a direct impact on business outcomes.
- Collaboration with data scientists, analysts, business stakeholders and external technology partners.
- Opportunities to solve complex data challenges and improve the performance, reliability and scalability of data platforms.
- Experience working across batch processing, workflow automation and real-time streaming solutions.
- The opportunity to contribute to data quality, platform efficiency and cost optimisation initiatives.
What skills you will learn
- Advanced development and optimisation of ETL and ELT pipelines on Google Cloud Platform.
- Cloud-based data processing using BigQuery, Cloud Storage, Dataflow, Cloud Composer, Pub/Sub and Cloud Functions.
- Workflow orchestration and automation using Apache Airflow and related technologies.
- Data pipeline monitoring, validation, cleansing and quality management.
- BigQuery query optimisation and cloud storage cost management.
- Kafka streaming deployment and integration using Python.
- Cross-functional collaboration and the translation of business requirements into scalable data solutions.
VOIS Equal Opportunity Employer Commitment
Join Us
Alert
Follow us on social media and #StayConnected