Overview:
Berkshire Hathaway HomeServices PenFed Realty (PenFed Realty), a wholly owned subsidiary of PenFed Credit Union (PenFed), is hiring an AI Product Manager. This fully remote opportunity is open to candidates located anywhere in the United States. The AI Product Manager owns the design, delivery, and adoption of PFR IT's AI agent products, beginning with the AI Personal Assistant (AIPA) program and expanding to the broader AI portfolio as it matures. This is a hands-on product role: the incumbent defines the problem, builds working prototypes, and ships production agents on the Gemini Enterprise Agent Platform alongside AI developers rather than handing specifications across a wall. The position is accountable for agent quality in production, including the evaluation sets that define acceptable performance, the automated pipeline that tests and releases changes, the escalation paths for when an agent is wrong, the consumption cost of the agents in service, and the measurable adoption of these tools by agents, sales managers, and administrative staff across the brokerage.
Responsibilities:
Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. This is not intended to be an all-inclusive list of job duties, and the position will perform other duties as assigned.
- AI Product Ownership: Own the roadmap for AIPA and its sub-agents end to end, including the prioritization decisions about what does not get built and the stakeholder conversations behind those decisions.
- Rapid Prototyping: Build working prototypes of proposed agent capabilities using AI-assisted development tools, delivering a functional demonstration rather than a specification document as the standard output of a discovery cycle.
- Agent Development: Build and ship agents directly using the Agent Development Kit (ADK) and Agent Studio, deploying to Agent Runtime and iterating on production behavior in partnership with AI developers.
- Evaluation Design: Build and own the evaluation sets that define acceptable performance for each agent, combining deterministic checks, model-graded rubrics, and human review, and grade the full agent trajectory, including tool selection, reasoning steps, and task completion, rather than the final response alone.
- Regression Testing and Release Automation: Treat evaluations as the test suite for non-deterministic software, running them automatically against every prompt revision, model upgrade, tool change, and data source change, and using the results as release gate criteria rather than deploying on subjective judgment.
- Continuous Evaluation: Monitor live agent quality in production, sample real traffic against evaluation criteria, and continually expand test sets with new edge cases so that offline results remain a reliable signal of production behavior.
- Automated Delivery Pipeline: Evaluate and recommend the AI development automation framework, then implement it, establishing an automated path from Jira backlog item through build, evaluation, and deployment, with Confluence maintained as the design and decision record.
- Token Cost Optimization: Monitor and report AI consumption across agents and development workflows, and optimize model selection, prompt design, context size, and evaluation run scope to reduce cost per interaction without degrading measured quality.
- Failure Mode Ownership: Design and document human-in-the-loop escalation, guardrails, and rollback procedures for agent errors before launch, and lead the response when an agent produces an incorrect result in front of a user.
- Regulatory and Compliance Alignment: Partner with Legal, Compliance, and brokerage leadership to ensure agents operate on contracts, transactions, and licensed activity meet regulatory requirements and preserve required human review.
- Discovery and Solution Selection: Evaluate proposed AI use cases against business value, data availability, and feasibility; determine whether a problem warrants an agent before development effort is committed; and assess Gemini Enterprise capabilities, foundation models available through Model Garden, and third-party tools to make build-versus-buy recommendations.
- Requirements Definition: Write requirements that developers can build from without a follow-up meeting, including data dependencies, tool and integration needs, latency expectations, and acceptance criteria.
- Data and Grounding: Identify, validate, and maintain the data sources and knowledge repositories that ground agent responses, and partner with data owners on quality, access, and retention.
- Platform Governance: Apply enterprise agent governance controls, including agent identity, registry, access, and observability, and shepherd new agents through information security and privacy review.
- Stakeholder Partnership: Work directly with sales leadership, branch management, operations, and administrative staff to translate frontline pain points into agent capabilities.
- Adoption and Outcomes: Drive measurable adoption of AI tools across the field in partnership with Learning and Development, and define and report the usage, quality, and business outcome metrics used to decide what to expand, rebuild, or retire.
- Applied AI in Daily Work: Use AI tools in the daily execution of the role, including research, analysis, drafting, and coding, and share effective practices with the broader PFR IT team.
Qualifications:
Equivalent combination of education and experience is considered.
- Bachelor’s degree in computer science, Information Systems, Business, or a related field.
- Minimum of three years of product management, technical program management, or software engineering experience, including at least one AI or machine learning powered feature shipped to production users and measured after launch.
- Hands-on development ability sufficient to independently prototype, build, and deploy AI agents, including working proficiency in Python and comfort operating in a source-controlled development workflow.
- Experience with an agent development framework such as the Agent Development Kit (ADK), LangGraph, CrewAI, or a comparable platform.
- Experience with Google Cloud Platform; familiarity with the Gemini Enterprise Agent Platform, Vertex AI, or Agent Engine preferred.
- Working fluency in AI evaluation, including the ability to design evaluation sets and grading criteria, distinguish deterministic, model-graded, and human evaluation methods and know when each applies, interpret results honestly, and explain why a capability that demonstrates well can still fail in production.
- Experience integrating automated evaluation into a release process, using tooling such as Agent Evaluation on the Gemini Enterprise Agent Platform, LangSmith, Phoenix, DeepEval, or a comparable framework.
- Experience building or operating continuous integration and deployment automation and connecting AI development workflows to Jira and Confluence.
- Familiarity with AI cost management, including token consumption monitoring, model selection tradeoffs, and prompt and context efficiency.
- Practical knowledge of prompt engineering, retrieval-augmented generation, tool calling, and agent orchestration patterns.
- Ability to query and analyze data independently using SQL or an equivalent analytics tool.
- Written communication is strong enough to withstand engineering scrutiny, producing requirements that generate specific questions about edge cases rather than broad confusion about intent.
- Experience delivering technology in a regulated industry preferred; real estate, mortgage, title, or financial services background is an advantage.
- Proficiency with Jira, Confluence, and GitHub.
- Demonstrated experience using AI tools in daily work is required.
Supervisory Responsibility
This position will not supervise employees.
Licenses and Certifications
Google Cloud Professional certifications preferred.
Work Environment
While performing the duties of this job, the employee is regularly exposed to a hybrid (in-office and remote) work setting.
*Most roles require working in an office setting with moderate noise and the ability to lift 25 pounds.*
Travel
Occasional travel to branch and operational locations.
Pay Transparency
The anticipated starting salary range for this role is $79,400.00 - $165,762.00
This position is eligible for an organizational performance based annual bonus, subject to board discretion and approval.
This position is eligible for an individual performance based annual bonus.
#LI-Hybrid