Trust & Safety Evaluator with English from the United States I 6-months project

TampereCompetitive0 applicants

About this role

Trust & Safety Evaluator conduct adversarial testing and safety evaluation of generative AI features.  Main tasks are crafting queries, evaluating the safety of generated content and providing critical feedback. This role requires creative thinking about potential misuse, deep cultural and linguistic knowledge, and the ability to identify subtle safety risks.

Project duration: 6 months

Responsibilities

  • Write, review and evaluate diverse and challenging queries designed to test the system's limits and expose problematic outputs. Queries will target specific risk topics including explicit and/or offensive content
  • Design and execute sequences of queries simulating realistic, unfolding conversations.
  • Craft attack scenarios using techniques like crescendo attacks and context manipulation to test the system
  • Age-Appropriate Safety Evaluation: to guide adversarial query crafting and safety evaluation.
  • Assign risk ratings to AI Generated content based on safety guidelines.
  • Detect and articulate potential biases in system-generated content
  • Identify subtle unsafe elements, inconsistencies, or problematic implications
  • Evaluate specific cultural and language knowledge across different cultural contexts.
  • Leverage familiarity with relevant domains, genres, cultural references, and industry context.
  • Complete adversarial testing tasks within time constraints
  • Maintain high accuracy and attention to detail while working at pace.
  • REQUIRED SKILLS
  • Language skills: English from United States as primary language is required
  • Deep cultural awareness of US English language: context, social norms, and values
  • Understanding of cultural references and industry context
  • Knowledge of how different cultural groups interpret content
  • Familiarity with behavioral patterns of younger audiences
  • Ability to evaluate content across different cultural contexts
  • Query crafting and scenario design
  • Techniques like crescendo attacks and context manipulation
  • Risk rating assignment bas

EU Requirements

Job Details

Posted4 October 2026
Closes3 November 2026

Contact

Similar Jobs

Finding similar jobs...