- We're looking for an AI Engineer to design and ship agentic workflows and LLM-powered features inside a U.S.-based SaaS platform serving the Non-Emergency Medical Transport (NEMT) industry across North America. You'll be the first dedicated AI engineer on the team — building dispatcher copilots, intelligent routing assistants, document and intake automation, and voice and chat agents for both customer-facing and internal operations use cases.
- This is not a research role, a data science role, or a general backend role that occasionally touches AI. AI features will be your full-time focus. You'll propose patterns, recommend tools, shape quality standards, and carry real influence over the direction the team takes — based on the strength of your thinking.
- The role is 100% Remoto, based in México, working in the EST time zone alongside a LatAm engineering hub, European-based engineers, and U.S.-based product and operations stakeholders.
- You have 3+ years of software engineering experience and at least 1 year of hands-on experience shipping LLM-powered features to real users in production — not prototypes or coursework.
- You bring sound architectural judgment on AI choices: you know when to use a third-party API, a self-hosted open-weight model, or no AI at all, and you make that call based on cost, latency, accuracy, and data residency constraints.
- You've built evaluation systems that prove AI features actually work — you can describe the metrics you defined, the test datasets you assembled, and how you detected regressions when models, prompts, or data changed.
- You communicate clearly in English in async, cross-timezone environments and can represent engineering constraints and AI feasibility to non-technical stakeholders.
- You're energized by being the person others come to when they ask "is this even an AI problem?"
- Architect and ship agentic workflows and LLM-powered product features including dispatcher copilots, routing assistants, document automation, and voice/chat agents — handling planning, memory, state management, tool orchestration, guardrails, and human-in-the-loop checkpoints.
- Design retrieval architectures for each use case, choosing between vector retrieval, long-context/cache-augmented, tool-based agentic retrieval, graph-based, or hybrid approaches based on the problem at hand.
- Build evaluation and observability systems that define accuracy, correctness, and safety metrics; assemble test datasets; run pre-release benchmarks; and monitor production quality to catch regressions.
- Harden AI features for production — managing APIs, latency and cost budgets, prompt versioning, observability, and PII/PHI safety alongside platform and product engineers.
- Build agentic pipelines for internal operations — automating marketing, sales, support, and back-office workflows where AI agents can meaningfully reduce manual effort.
- Partner with product and engineering leadership during feature ideation and scoping, providing pragmatic input on level of effort, technical feasibility, and the realistic likelihood that proposed AI approaches will work in production.
- Educate and enable the broader engineering team to incorporate agentic flows into their own product work so AI capabilities are woven across the team, not siloed.
- Champion AI-assisted development practices across the LatAm Hub engineering team.
- A production-first mindset — you don't ship LLM features you can't measure, and you build the scaffolding to prove quality before launch.
- Architectural judgment that goes beyond picking the trendiest tool — you choose based on constraints, not hype.
- Strong async written communication and comfort collaborating across Mexico, Europe, and U.S. time zones.
- A team-enablement instinct — you share what you know and raise the AI maturity of the engineers around you.
- 3+ years of software engineering experience with 1+ year shipping production LLM/AI features to real users.
- Strong Python skills and comfort with TypeScript.
- Hands-on experience with at least one agent framework (LangChain/LangGraph, LlamaIndex, OpenAI Agents SDK, or equivalent).
- Production experience with major LLM provider APIs (Anthropic Claude, OpenAI, or AWS Bedrock).
- Experience with multiple retrieval and context-augmentation approaches and the judgment to choose between them.
- Demonstrated ability to define quality metrics, build test datasets, and detect regressions in production AI systems.
- Working professional English with strong async written communication.
- Strongly Preferred:
- Voice AI experience (ElevenLabs, VAPI, LiveKit, Deepgram, or equivalent).
- Experience in a healthcare or regulated-data context (HIPAA, PII/PHI handling, audit logging, data minimization).
- Local/self-hosted LLM experience running open-weight models on-prem or in a VPC.
- Anthropic ecosystem familiarity — Claude API, extended thinking, prompt caching, tool use, computer use, or Agents SDK.
- LLM observability and eval tooling experience (LangSmith, Braintrust, Phoenix, Helicone, or similar).
- Cost and latency optimization at LLM scale: prompt caching, model routing, token budgeting.
- Traditional ML or data science background.
- Django and PostgreSQL experience.
- Multi-tenant SaaS experience.
- Open-source AI contributions or public agent projects.
Skills
OpenAI API
Amazon Web Services (AWS)
Python