Skip to content
← Back to job listings

AI Architect (Voice AI)

neurons-lab.com · Poland

RemoteExternal listingcontractabout 16 hours ago

About The Role

About the project (description, duration, stage)The client is the largest US network of in-home veterinary hospice and end-of-life care. A major US private-equity sponsor drives the AI program and plans more projects across its portfolio.We built a real-time voice copilot for their Veterinary Care Coordinators (VCCs). The copilot listens to live calls with pet families. It extracts appointment and clinical fields while the call runs. It fills the client's scheduling system through a Chrome extension. A second workstream, the Vet Visit Copilot, sends each vet an AI pre-visit briefing by email (Amazon SES).Next is the production phase. Stage: production SOW in executive alignment; start expected September 2026.Duration: multi-month, with strong extension probability. 0.5 FTE minimum; ramp toward 1.0 FTE as production scales.Why the role is open: the current architect moves to another strategic build. He stays at 0.15–0.2 FTE for supervision and knowledge transfer during ramp-up, so the new architect gets a structured handover.ObjectiveOwn the technical architecture and delivery of the voice copilot from validated PoC to productionHit the bar this client tests against: latency, accuracy, concurrency, and costKeep expectations aligned: production polish is in scope now; protect the team from silent scope creepTransfer knowledge continuously to the client's team and Neurons Lab engineersAreas of ResponsibilityTechnical architecture & hands-on implementationOwn the full pipeline: streaming speech-to-text, LLM field extraction, Chrome-extension delivery, and AWS infrastructureDrive latency work: cut P95 from ~6s toward ~2s; remove post-processing corner cases (occasional ~1min lag on one field type)Run model A/B tests (current pair: Claude Haiku vs GPT Luna) with golden-set evaluation for phonetic name and email accuracyOwn evaluation and cost: Langfuse traces, accuracy dashboards, real per-call cost from live calls, and an optimization planHarden for production: 5–10+ concurrent calls, strict data isolation between users, monitoring, alerting, and safe rollbackShip epics end to end (example: the SES email briefing service); always keep a demo fallback so a live session never failsWorking with client stakeholdersFront technical discussions with a meticulous client; VCCs test edge cases and expect production qualityPresent concrete system behavior, with numbers — this account rewards evidence, not slidesHold the scope line: tie every feedback item to the SOW; route roadmap items (learning loop, persistent memory) to future phasesKeep internal discussions internal; all client-facing materials pass ADM review before sendingTeam & knowledgeLead the AI Engineer and the pod: set tasks, review output, unblock fastAbsorb the handover from the outgoing architect (0.15–0.2 FTE supervision window) and become independent fastRun knowledge-transfer sessions; the project must have no single point of failureSupport the production SOW with estimates and architecture options when the account team asksSkillsReal-time voice pipelines: streaming STT, turn handling, low-latency LLM inference — hands-onLLM engineering: prompt engineering, structured extraction, guardrails, model A/B evaluationObservability and evals: Langfuse or similar; golden datasets; latency, accuracy, and cost dashboardsAWS: Bedrock, serverless patterns, SES; token economics and per-call cost engineeringFull-stack pragmatism: strong Python; enough TypeScript / Chrome-extension knowledge to own the integrationClear spoken and written English for demanding US executivesKnowledgeContact-center / agent-assist patterns and metrics (handle time, cost per call, concurrency)Production LLM operations: load testing, data isolation, incident handlingNice to have: empathy-sensitive domains (healthcare, veterinary, insurance) and PE-sponsored rolloutsExperienceKey characteristics (screen for all four):Voice AI in production — mandatory. Shipped at least one real-time voice or speech product to real users (agent assist, voice bot, live transcription copilot). Candidates will demo real artifacts at the interview.6+ years hands-on AI/ML engineering, with strong recent LLM production practiceLatency and reliability record. Can show measured P95 reductions and concurrency fixes on a live systemConsulting / client-facing seniority. Calm and precise under detailed UAT scrutiny; manages expectations wellNice to have:Chrome extension delivery; telephony / streaming stacks (Amazon Connect, Twilio, LiveKit)Langfuse in productionUS client experience with Eastern-time overlap

This is an external listing. JobSpring does not represent or verify the employer. Report this listing