Job description
Responsibilities:
• Support development and optimisation of foundation model training workflows.
• Work on transformer architectures, model fine-tuning, and evaluation systems.
• Build and maintain data intelligence and curation pipelines.
• Implement data preprocessing, deduplication, tokenisation, and quality scoring workflows.
• Support RLHF/alignment and benchmarking systems.
• Work with scientific and multilingual datasets.
• Collaborate with AI infrastructure and platform engineering teams.
• Contribute towards scalable GPU/TPU-based AI systems.
Requirements:
• Strong Python programming skills.
• Understanding of Machine Learning and Deep Learning fundamentals.
• Hands-on exposure to PyTorch or similar ML frameworks.
• Understanding of transformers, LLMs, or foundation model concepts.
• Familiarity with NLP preprocessing or data pipeline concepts.
• Strong debugging and problem-solving mindset.
• Comfortable working in Linux environments