LLM Model Response Evaluation
TLDR
Evaluate AI-generated text, images, audio, video, HTML widgets, and PDFs using trusted research and defined quality rubrics.
- Evaluating UI widgets, infographics, image factuality, side-by-side comparisons, and similar AI evaluation activities.
- The work may involve text, images, audio, video, HTML widgets, PDFs, or combinations of these modalities.
- The work is domain-agnostic and may cover topics across arts, culture, history, science, engineering, and more.
- Resources will be expected to independently research unfamiliar topics using trusted sources before making evaluation decisions.
- Each task will include detailed project guidelines within the evaluation platform.
- 3+ years of hands-on experience in LLM / GenAI data evaluation.
- Master's or PhD required (PhD candidates strongly preferred).
- Ability to research unfamiliar topics using trusted sources and make well-supported judgments.
- Comfortable evaluating content across multiple modalities
Flexible and remote work
Variable workload: Accept or decline tasks based on your availability
No guaranteed hours: Workload may vary weekly
Benefits
Flexible Work Hours
Accept or decline tasks based on your availability
Remote-Friendly
Flexible and remote work
Lifted, an Upwork Company™, connects enterprise clients with skilled freelance professionals to support specialized projects in safety, finance, and artificial intelligence. By leveraging a network of experts, we streamline the process of sourcing talent for complex tasks, ensuring high-quality results that drive innovation and efficiency in high-demand sectors.