Data Engineer - RAG & Generative AI:
- Type of Employment:Contractor
- Work Modality:100% Remote
- Work Schedule:Full-time
- Location:LATAM
- Start Date:August 3, 2026
- End Date:December 31, 2026 (with the possibility of extension)
About Our Client:
Our client is a fast-growing technology company building AI-powered products that help organizations unlock the full potential of their data. They foster a collaborative, engineering-driven culture where Data, AI, and Product teams work closely together to build scalable, production-ready solutions using modern cloud technologies.
About the Role:
In this role, you'll transform raw, unstructured data into AI-ready knowledge by developing scalable ingestion, processing, embedding, and vector indexing pipelines. You'll work closely with AI Engineers, Machine Learning teams, and Product stakeholders to ensure our AI systems have fast, reliable access to high-quality data that drives accurate and relevant responses.
What You'll Do:
- Process and transform data from multiple formats, including PDFs, Word documents, HTML, JSON, and XML.
- Implement document parsing, chunking, embedding generation, and vector indexing workflows.
- Manage, optimize, and maintain vector databases for semantic search and retrieval.
- Build data validation, metadata enrichment, and deduplication processes.
- Monitor, troubleshoot, and optimize production data pipelines for reliability and performance.
- Collaborate with AI Engineers and Product teams to continuously improve retrieval quality and system performance.
What We're Looking For:
- 3+ years of experience in Data Engineering or building production data pipelines.
- Strong proficiency in Python and SQL.
- Experience working with cloud platforms such as AWS, Azure, or GCP.
- Hands-on experience with data engineering tools like Spark, Airflow, Kafka, or similar technologies.
- Experience with vector databases such as Pinecone, Weaviate, Milvus, Chroma, or Qdrant.
- Familiarity with documen
📌 Data Engineer (Lima)
🏢 ScaleUp
📍 Lima