LLM Model Response Evaluation
Lifted (an Upwork Company) · United States · 2026-10-01
About this role
• Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
• The work may involve text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
• The work is domain-agnostic and may cover topics across arts, culture, history, science, engineering, and more.
• Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions.
• Each task will include detailed project guidelines within the evaluation platform.
• 3+ years of hands-on experience in LLM / GenAI data evaluation.
• Bachelor's Degree required
• Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
• Comfortable evaluating content across multiple modalities
Flexible and remote work
Variable workload: Accept or decline tasks based on your availability
No guaranteed hours: Workload may vary weekly
Our client, a global technology company that helps businesses build, train, and manage AI systems is looking for experts to evaluate model-generated content against defined quality rubrics such as factuality, consistency, aesthetics, and other evaluation criteria.
Originally posted on Himalayas
Skills asked for
- llm
- html
Similar jobs
- Senior Applied Scientist, Efficient LLM Inference & Model OptimizationNebius · Amsterdam
- DACH Thesis: Using AI (LLM) in Process ModelingBroadpin · Ettlingen
- Senior Applied Scientist, Efficient LLM Inference & Model OptimizationNebius · Palo Alto
- Machine Learning Engineer, Model Evaluations (Speech LLM) - San FranciscoPlaud · San Francisco
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.