EXTERNAL LISTING · TURING

Software Engineer – AI Code Evaluation & Benchmarking (US candidates only)

Remote contractor opportunity for experienced Software Engineers to evaluate AI-generated code, benchmark coding capabilities, debug solutions, and contribute to datasets and evaluation frameworks for advanced AI models.

Software Engineering / AI EvaluationRemote

About this opportunity

Turing is seeking experienced Software Engineers to evaluate, benchmark, and improve the coding capabilities of frontier AI models. The role focuses on reviewing AI-generated code, validating solutions against real-world software engineering tasks, identifying correctness and quality issues, debugging implementations, and contributing to high-quality evaluation datasets, benchmarks, and grading rubrics.

What the listing describes

  • Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements
  • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes
  • Debug code, reproduce issues, and verify fixes across different programming environments
  • Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy
  • Create, refine, and maintain evaluation datasets for coding tasks
  • Create and maintain coding benchmarks and grading rubrics
  • Identify edge cases and failure modes in AI-generated software solutions
  • Identify areas where AI systems struggle with software engineering problems
  • Document findings clearly and provide structured feedback
  • Contribute to evaluation quality and consistency
  • Collaborate with project teams to establish quality standards and evaluation methodologies

Requirements shown on the source

  • Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related technical field
  • 3+ years of professional software engineering experience
  • Strong proficiency in one or more of Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL
  • Strong understanding of data structures and algorithms
  • Strong understanding of software design principles
  • Strong understanding of debugging methodologies
  • Experience performing code reviews and evaluating code quality in production or large-scale codebases
  • Ability to analyze complex technical problems and assess solution correctness with minimal supervision
  • Familiarity with version control systems such as Git
  • Familiarity with modern software development workflows
  • Strong written communication skills
  • Strong attention to detail
  • Availability for at least 4 hours per day and a minimum of 20 hours per week
  • Ability to provide 4 hours of overlap with PST

Relevant backgrounds mentioned

  • Experience with AI/ML data annotation
  • Experience with natural language processing
  • Experience with prompt engineering
  • Experience with AI model evaluation
  • Experience with LLM-related projects
  • Experience evaluating AI-generated code
  • Experience creating software engineering benchmarks
  • Experience with software quality assessment
  • Experience working with Python and Docker

Details to verify before applying

Country eligibility
United States
Compensation
No pay rate was stated on the source page reviewed.
Last checked

Source and current details

Our summary is not the full job description. Review the external listing for the latest requirements, terms and application process.

Open the source listing on Turing