Job Overview
- Design and build robust data ingestion pipelines, transforming raw CSV/TSV data into structured, queryable datasets.
- Model custom entities and map diverse data sources into standardised target schemas.
- Build and manage data environments using SQLite, Google Cloud SQL and BigQuery.
- Develop and maintain Dockerised services for data loading, embedding generation and data-serving applications.
- Deploy and operate production workloads on Google Cloud Run, with a focus on scalability, reliability and high availability.
- Work across Redis, VPC networking, Artifact Registry and Cloud Storage.
- Implement secure cloud architectures using Cloud IAM, VPC and Cloud Run ingress controls.
- Automate data workflows and deployments through GitHub Actions, including scheduled ingestion, ID mapping, embedding generation and container builds.
- Build and integrate AI-powered capabilities, including natural-language querying, semantic search, embeddings and LLM-powered workflows.
- Contribute to emerging agentic architectures and integrations, including technologies such as Model Context Protocol (MCP).
Ready to Apply?
Take the next step in your career journey
Stand out with a professional resume tailored for this role