Applied AI Researcher

Kasmo Inc

  • Jersey City, NJ
  • 15 days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Web Services (AWS)unmatched
    • Artificial Intelligence (AI)unmatched
    • Bank Managementunmatched
    • Banking Servicesunmatched
    • Benchmarkingunmatched
    • Business Caseunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Computational Linguisticsunmatched
    • Computer Scienceunmatched
    • Data Analysisunmatched
    • Data Recoveryunmatched
    • Data Setsunmatched
    • Develop Methodologiesunmatched
    • Documentationunmatched
    • Experiment Designunmatched
    • Financial Complianceunmatched
    • Information Retrievalunmatched
    • Know Your Customer (KYC)unmatched
    • Legalunmatched
    • Machine Toolunmatched
    • Mathematicsunmatched
    • Metricsunmatched
    • Model Validationunmatched
    • Natural Language Processing (NLP)unmatched
    • Open Sourceunmatched
    • Patentsunmatched
    • Performance Modelingunmatched
    • Prototypingunmatched
    • Publicationsunmatched
    • Research Skillsunmatched
    • Riskunmatched
    • Safety/Work Safetyunmatched
    • Scientific Researchunmatched
    • Statisticsunmatched
    • Underwritingunmatched
    • Use Casesunmatched

    Description



    Description:
    This role requires working onsite 4 days per week, and a F2F interview at the client s Jersey City location is mandatory.
    Only USCs/GCs are eligible
    Applied AI Researcher
    GenAI / NLP / Agentic AI / Applied Machine Learning
    Level
    Research-focused Individual Contributor
    Target / alternate titles
    Applied Scientist; Research Scientist - NLP; AI Research Scientist; ML Researcher; NLP Scientist; GenAI Researcher; Applied ML Scientist
    Core keywords
    applied research, LLM, NLP, transformers, RAG, retrieval, model evaluation, experiments, robustness, embeddings, synthetic data, multimodal, PyTorch, Hugging Face, banking AI, AWS AI, AIRP
    Recruiter red flags
    Academic-only profile with no applied delivery; weak experimental design; cannot translate research into AIRP-ready business or engineering requirements; no awareness of regulated data constraints.
    Role purpose
    Bridge advanced AI research and practical enterprise use cases by validating models, methods, and prototypes that can become production-grade AIRP solutions. The role focuses on measurable business value, rigorous experimentation, model behavior, and safe translation of research into banking-relevant applications.
    Client-specific emphasis
    Research must be grounded in enterprise business use cases, not generic AI experimentation.
    Candidates should understand how model, retrieval, data, evaluation, latency, cost, and safety decisions affect production delivery on AIRP.
    Cloud/AWS awareness is valuable because successful research outputs must be handed off to engineering teams building on AWS-hosted AIRP.
    Primary ownership
    Applied research agenda for LLMs, NLP, RAG, evaluation, multimodal AI, and agentic workflows relevant to enterprise use cases.
    Prototypes, experiments, benchmark design, model-selection recommendations, and production-readiness evidence.
    Research-to-production handoff with AI engineering, AIRP platform, product, risk, and governance teams.
    Key responsibilities
    Conduct applied research in LLMs, GenAI, NLP, information retrieval, multimodal AI, synthetic data, and agentic AI.
    Design experiments to evaluate model performance, robustness, safety, scalability, interpretability, enterprise usefulness, and production feasibility.
    Prototype AI solutions for KYC, credit underwriting, governance tracking, pitch book generation, Banker 360, Customer 360, deal library intelligence, financial crime quality, and sanctions screening.
    Develop evaluation methodologies using golden datasets, adversarial testing, offline benchmarks, human review, business outcome metrics, and risk-specific acceptance criteria.
    Assess prompt optimization, RAG, fine-tuning, instruction tuning, synthetic data generation, distillation, and model adaptation techniques.
    Document model limitations, data assumptions, hallucination patterns, bias risks, performance boundaries, and control recommendations for regulated deployment.
    Collaborate with engineers to convert prototypes into production-ready AIRP requirements, including latency, cost, observability, security, and AWS/cloud deployment considerations.
    Track emerging AI research and translate relevant advances into practical recommendations for the enterprise.
    Must-have candidate profile
    Advanced degree preferred, usually MS or PhD in AI, ML, computer science, statistics, computational linguistics, mathematics, or related field.
    Strong foundation in machine learning, deep learning, NLP, transformers, information retrieval, and generative AI.
    Hands-on experience with LLMs, embeddings, RAG, model evaluation, and applied GenAI experimentation.
    Python skills with PyTorch, TensorFlow, Hugging Face, scikit-learn, or equivalent research frameworks.
    Ability to design rigorous experiments and communicate findings to technical, product, business, risk, and governance stakeholders.
    Ability to translate research results into production requirements suitable for an AWS-hosted enterprise platform.
    Preferred experience
    Research or applied science experience in banking, finance, compliance, risk, legal, operations, financial crime, sanctions, or enterprise knowledge systems.
    Experience with AWS Bedrock, SageMaker, vector search, MLflow, Databricks, model evaluation tooling, or cloud-based experimentation environments.
    Publications, patents, internal research contributions, open-source AI contributions, or prior research-to-production handoffs.
    Familiarity with Responsible AI, model validation, privacy constraints, audit documentation, and regulated deployment environments.
    Initial screening questions
    What research idea did you convert into a prototype or production capability?
    How would you design an evaluation harness for an LLM-based banking use case such as KYC, underwriting, or sanctions screening?
    How do you determine whether fine-tuning, RAG, prompting, or model adaptation is the right approach?
    How do you account for latency, cost, safety, and AWS/cloud deployment constraints in applied research?
    What failure modes did you Client and how did you mitigate them?
    How do you communicate model limitations to non-research stakeholders?

    Numbers & Facts

    LocationJersey City, NJ

    Similar Jobs

    See more jobs