Remote | Machine Learning & NLP Research Specialist — $75–$105/hour
Job Description
This role supports an advanced AI initiative focused on identifying reasoning and capability gaps in frontier models. Selected professionals will design challenging real-world ML and NLP tasks, develop executable reference solutions, evaluate model performance, and analyse failures across language understanding, generation, retrieval, training workflows, and applied machine learning systems.
Key Responsibilities
ML & NLP Task Development
• Design challenging machine learning and natural language processing problems based on practical research or industry experience
• Create tasks involving model training, evaluation, language understanding, generation, retrieval, or applied ML pipelines
• Target specific reasoning, implementation, and capability gaps in advanced AI models
• Define clear specifications, expected behaviour, constraints, datasets, and evaluation criteria
Reference Solutions & Python Development
• Develop accurate reference solutions and supporting materials using Python
• Integrate tasks into agent-based development and evaluation environments
• Create executable tests, validation scripts, scoring logic, or model pipelines where appropriate
• Ensure reference implementations are technically sound, reproducible, and appropriately challenging
Model Evaluation & Failure Analysis
• Evaluate model and agent performance across assigned ML and NLP tasks
• Compare generated outputs with reference solutions and expected results
• Identify tasks where models demonstrate meaningful limitations or inconsistent behaviour
• Classify failures involving reasoning, implementation, retrieval, language understanding, generation, or instruction adherence
• Document findings through clear and technically detailed written analysis
Quality Calibration & Collaboration
• Review tasks and evaluation methods developed by other machine learning specialists
• Maintain consistent standards for difficulty, accuracy, realism, and technical quality
• Participate in calibration and peer-review activities
• Collaborate with other subject matter experts to improve evaluation coverage and reliability
Ideal Profile
Strong candidates may have:
• Deep hands-on experience in machine learning, natural language processing, or both
• Practical proficiency in Python demonstrated through professional, academic, or open-source work
• Strong understanding of modern ML and NLP methods, including transformers and large language models
• Experience with model training, fine-tuning, evaluation, retrieval, or production ML pipelines
• Familiarity with PyTorch, TensorFlow, JAX, Hugging Face, or comparable frameworks and tooling
• Ability to design realistic technical problems and develop complete reference solutions
• Strong written communication and the ability to explain complex model behaviour clearly
• Reliable availability for approximately 20 hours per week
Educational Background
• A degree in computer science, machine learning, artificial intelligence, computational linguistics, data science, or a related technical field is highly relevant
• Graduate or doctoral research in machine learning, NLP, language modelling, information retrieval, or related areas may be especially valuable
• Equivalent professional experience in applied ML or NLP may also be considered
• Published research, open-source contributions, or production ML work may strengthen an application
Nice to Have
• Experience with transformer architectures, large language models, fine-tuning, or alignment methods
• Background in information retrieval, embeddings, reranking, search, or retrieval-augmented generation
• Experience with text classification, sequence modelling, summarisation, translation, or language generation
• Familiarity with benchmark design, model evaluation, error analysis, or adversarial testing
• Experience building executable technical assessments or automated evaluation pipelines
• Previous involvement in AI training, model evaluation, data annotation, or quality-review programmes
• Familiarity with agent-based development environments and structured technical rubrics
Why This Opportunity
• Apply advanced machine learning and NLP expertise to frontier AI evaluation
• Design realistic tasks grounded in practical research and engineering experience
• Analyse how modern models perform across language understanding, generation, retrieval, and applied ML workflows
• Help identify meaningful capability gaps and opportunities for model improvement
• Work remotely with a focused part-time commitment and competitive hourly compensation
Contract Details
• Part-time W-2 contingent employment arrangement
• Fully remote role available to candidates based in the United States
• Expected commitment of approximately 20 hours per week
• Competitive rates between $75–$105 per hour depending on expertise and project scope
• Working proficiency in Python is required
• Work may include onboarding, technical calibration, task development, model evaluation, and peer review
• Project scope, workload, and duration may be adjusted according to programme requirements and performance
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.
