Rose International logo

AI Red Team Analyst - LLM Jailbreaking & Adversarial Testing

Rose International
  • Austin, Texas
  • $45–$50 Per Hour
  • Full-time
4 days ago

Job Description

AI Red Teamer

Job Description Client Location: Onsite 5 days a week in Austin, TX (downtown; parking fees not reimbursed); open to Onsite in New York, NY (Hudson Yards) Potential exposure to graphic or objectionable content (imagery, video, or written)

Required Skills:

  • 5+ years of experience designing / creating adversarial prompts and model jailbreaking
  • Creative & narrative construction — ability to build compelling multi-turn scenarios that exploit linguistic nuance
  • Psychological insight — understanding of human vulnerabilities, manipulation tactics, and social engineering principles
  • Adversarial/attack mindset — systematic ability to identify and exploit failure modes in AI systems
  • Resilience under exposure to graphic/objectionable content - sustained capacity to engage with disturbing material professionally

Nice to Have Skills:

  • Prior professional red teaming experience (AI or traditional security)
  • Experience with LLMs and generative AI products (ChatGPT, Claude, Gemini, Llama)
  • Prompt engineering techniques and jailbreak methodologies
  • Basic encoding/obfuscation knowledge (Base64, ROT13, Unicode tricks)
  • Scripting ability (Python or similar) to complement creative attacks
  • Understanding of AI safety concepts: RLHF, alignment, model evaluation frameworks
  • Socio-technical risk analysis experience
  • Knowledge of internet subcultures and emerging online threats
  • Experience with data annotation pipelines or labeling tools

Day to Day:

  • Design and execute novel, multi-turn adversarial attacks (emotional manipulation, roleplay, social engineering, authority exploitation) to bypass model safeguards
  • Evaluate model outputs for actual harm and real-world risk using user risk taxonomy
  • Probe for agentic vulnerabilities (privilege escalation, indirect prompt injection, scope creep in multi-authority systems)
  • Annotate model failures, classify vulnerabilities, and produce reproducible adversarial test cases
  • Collaborate with AI researchers, engineers, and domain experts to translate findings into actionable improvements
  • Contribute to refining red teaming taxonomy, benchmarks, and tooling infrastructure
  • Test across multiple modalities: text, image, audio, and video
  • Stay current with evolving adversarial techniques, internet subcultures, and AI safety research

About the Role:

We are seeking creative, resilient, and highly motivated AI Red Teamers to join our Red Teaming team. In this role, you will be at the forefront of AI safety, identifying and mitigating risks in our advanced language models. Unlike automated testing, which handles baseline coverage and known attack patterns efficiently, this role focuses on the ~40% of vulnerabilities that require human ingenuity, psychological insight, and creative out-of-distribution thinking. You will interact with our models across text, image, audio, and video modalities to uncover weaknesses, evaluate actual output harm, and stress-test our systems against novel adversarial attacks. This is intense, high-impact work at the frontier of AI safety — and it offers direct influence on how Client AI systems behave in the real world.

What You'll Do:

  • Creative Adversarial Testing
  • Design and execute novel, multi-turn adversarial attacks — including emotional manipulation, roleplay, social engineering, and authority exploitation — to bypass model safeguards and surface harmful capabilities.
  • Vulnerability Assessment
  • Evaluate model outputs for actual harm and real-world risk, not just policy violations. Apply Client user risk taxonomy to prioritize testing across user types, from casual users to agentic systems.
  • Agentic and Emerging Threat Testing
  • Probe for agentic vulnerabilities such as privilege escalation, indirect prompt injection, and scope creep in multi-authority systems — the next frontier of AI risk.
  • Data Annotation and Reporting
  • Generate high-quality human evaluation data by annotating model failures, classifying vulnerabilities, and producing reproducible adversarial test cases that engineering and safety teams can act upon.
  • Cross-Functional Collaboration
  • Partner with AI researchers, engineers, and domain experts to translate findings into actionable improvements. Contribute to refining our red teaming taxonomy, benchmarks, and tooling infrastructure.
  • Continuous Learning
  • Stay current with evolving adversarial techniques, internet subcultures, and AI safety research to continuously sharpen your attack strategies

What We're Looking For:

We actively seek candidates from non-traditional backgrounds. Our most effective red teamers come from creative, humanities, mental health, and special education fields — not exclusively from technical roles. Minimum Qualifications:

  • Creative and Psychological Insight: Background in creative writing, humanities, mental health counseling, psychology, or special education — with a demonstrated ability to construct compelling narratives, exploit linguistic nuance, and identify psychological vulnerabilities.
  • Adversarial Mindset: A natural inclination to think like an attacker and push systems to their limits. You should find genuine satisfaction in discovering unexpected failure modes.
  • Adaptability: Comfort switching between modalities (text, image, audio, video) and rapidly adjusting to new model behaviors, testing priorities, and task types.
  • Communication Skills: Excellent written and verbal communication skills, with the ability to document and explain complex vulnerabilities clearly to both technical and non-technical audiences.
  • Resilience and Balance: Capacity to sustain well-being while engaging in psychologically demanding work. This role involves regular exposure to graphic and objectionable content, including violence, exploitation, and self-harm scenarios. Comprehensive wellness support is provided.

Important Notice:

This role involves exposure to graphic and/or objectionable content, including but not limited to graphic images, videos, audio, and writings; offensive or derogatory language; and other potentially disturbing material such as child exploitation, graphic violence, self-injury, and animal abuse. Testing may also require verbalizing model prompts containing references to such content. Wellness infrastructure and opt-out policies are in place to support all team members.

Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Los Angeles Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, qualified applicants will be considered for assignment with arrest and conviction records. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, meet client expectations, standards, and accompanying requirements, and safeguard business operations and company reputation.

Numbers & Facts

LocationAustin, Texas
Job TypeFull-time
IndustryStaffing/Employment Agencies
Salary$45–$50 Per Hour
Company Size2,500 to 4,999 employees
Websitehttps://www.roseint.com/

About Company

Founded in 1993 by Sue Bhatia, Rose International is one of the nation's leading minority- and woman-owned providers of Staffing and Total Talent Solutions. We serve companies in all 50 states and employ thousands of people across the country.

Skills

  • Analysis Skillsunmatched
  • Artificial Intelligence (AI)unmatched
  • Audiovisualunmatched
  • Benchmarkingunmatched
  • Business Operationsunmatched
  • Communication Skillsunmatched
  • Computer Securityunmatched
  • Constructionunmatched
  • Creative Writingunmatched
  • Cross-Functionalunmatched
  • Customer/Consumer Behaviorunmatched
  • Data Managementunmatched
  • Data Modelingunmatched
  • Establish Prioritiesunmatched
  • Injectionsunmatched
  • Internet Securityunmatched
  • Machine Toolunmatched
  • Modeling Languagesunmatched
  • Presentation/Verbal Skillsunmatched
  • Psychiatry and Mental Healthunmatched
  • Psychologyunmatched
  • Python Programming/Scripting Languageunmatched
  • Riskunmatched
  • Risk Analysisunmatched
  • Risk Managementunmatched
  • Scripting (Scripting Languages)unmatched
  • Security Attacksunmatched
  • Social Engineeringunmatched
  • Special Educationunmatched
  • Stress Testingunmatched
  • Surface Modelingunmatched
  • Taxonomiesunmatched
  • Test Automationunmatched
  • Test Caseunmatched
  • Training Data Setsunmatched
  • Unicodeunmatched
  • Writing Skillsunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder