{"schemaVersion":"jobsearcher.job.v1","id":"abd103cd6da29b1ab77e3003","url":"https://jobsearcher.com/jobs/abd103cd6da29b1ab77e3003","canonicalUrl":"https://jobsearcher.com/jobs/abd103cd6da29b1ab77e3003","title":"AI Platform Engineer","description":"Job Title: Data Engineer / AI Engineer (Agentic AI Platform Financial Data)\nLocation: Philadelphia, PA (Hybrid)\nDuration: 12+ months contract\nAbout the Role: We are building a platform that converts unstructured financial data (emails, corporate actions, index announcements) into high-quality, structured datasets used by financial institutions. This is not a typical \"LLM wrapper\" role. You will work on systems that:\nExtract data from noisy, inconsistent sources\nValidate and reconcile outputs across multiple inputs\nEnsure correctness, traceability, and auditability\nThe challenge is not just applying LLMs-it's making them reliable in production for financial workflows.\nWhat You'll Work On\nDesigning pipelines that process high-volume financial documents (batch + near real-time)\nBuilding LLM-powered extraction workflows (classification, parsing, summarization)\nImplementing validation layers (rule-based + model-based) to reduce hallucinations\nDeveloping retrieval systems using embeddings and vector search\nArchitecting end-to-end systems: ingestion processing storage serving\nEnsuring data quality, observability, and fault tolerance\nCollaborating with product to turn messy data into usable financial intelligence\nCore Requirements\nStrong Python and backend/data engineering experience\nExperience building production data pipelines (ETL, streaming, or async systems)\nSolid understanding of distributed systems and failure modes\nExperience working with LLM-based systems in production: Prompt design, Output validation, Retry/fallback strategies, Evaluation and monitoring\nExperience with data storage systems (SQL + NoSQL)\nFamiliarity with cloud infrastructure (AWS or similar)\nPreferred Experience\nExperience with RAG / vector search systems\nBackground in financial data or capital markets\nExperience with streaming systems (Kafka, etc.)\nExperience building multi-step or agent-style workflows\nWhat Makes This Role Interesting\nWork on high-accuracy AI systems where correctness matters\nSolve real problems around: LLM reliability and hallucination mitigation\nData consistency across conflicting sources\nReal-time vs correctness tradeoffs\nBuild systems used in financial decision-making workflows\nHigh ownership over core architecture in an early-stage environment\nNice to Know (but not required)\nExperience with orchestration tools (Airflow, etc.)\nExposure to evaluation frameworks for LLMs\nExperience working with large-scale document processing\nTech Stack (Representative, not exhaustive)\nPython, APIs, async processing\nLLM APIs + embeddings\nSQL / NoSQL databases\nCloud infrastructure (AWS)\nData pipelines and streaming systems\nVector Databases\nFor applications and inquiries, contact: hirings@openkyber.com","company":"Openkyber","rawCompany":"openkyber","city":"Alaska","state":"MI","isRemote":false,"isActive":false,"createdAt":"2026-08-04T16:44:26.798Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Platform Engineer","description":"Job Title: Data Engineer / AI Engineer (Agentic AI Platform Financial Data)\nLocation: Philadelphia, PA (Hybrid)\nDuration: 12+ months contract\nAbout the Role: We are building a platform that converts unstructured financial data (emails, corporate actions, index announcements) into high-quality, structured datasets used by financial institutions. This is not a typical \"LLM wrapper\" role. You will work on systems that:\nExtract data from noisy, inconsistent sources\nValidate and reconcile outputs across multiple inputs\nEnsure correctness, traceability, and auditability\nThe challenge is not just applying LLMs-it's making them reliable in production for financial workflows.\nWhat You'll Work On\nDesigning pipelines that process high-volume financial documents (batch + near real-time)\nBuilding LLM-powered extraction workflows (classification, parsing, summarization)\nImplementing validation layers (rule-based + model-based) to reduce hallucinations\nDeveloping retrieval systems using embeddings and vector search\nArchitecting end-to-end systems: ingestion processing storage serving\nEnsuring data quality, observability, and fault tolerance\nCollaborating with product to turn messy data into usable financial intelligence\nCore Requirements\nStrong Python and backend/data engineering experience\nExperience building production data pipelines (ETL, streaming, or async systems)\nSolid understanding of distributed systems and failure modes\nExperience working with LLM-based systems in production: Prompt design, Output validation, Retry/fallback strategies, Evaluation and monitoring\nExperience with data storage systems (SQL + NoSQL)\nFamiliarity with cloud infrastructure (AWS or similar)\nPreferred Experience\nExperience with RAG / vector search systems\nBackground in financial data or capital markets\nExperience with streaming systems (Kafka, etc.)\nExperience building multi-step or agent-style workflows\nWhat Makes This Role Interesting\nWork on high-accuracy AI systems where correctness matters\nSolve real problems around: LLM reliability and hallucination mitigation\nData consistency across conflicting sources\nReal-time vs correctness tradeoffs\nBuild systems used in financial decision-making workflows\nHigh ownership over core architecture in an early-stage environment\nNice to Know (but not required)\nExperience with orchestration tools (Airflow, etc.)\nExposure to evaluation frameworks for LLMs\nExperience working with large-scale document processing\nTech Stack (Representative, not exhaustive)\nPython, APIs, async processing\nLLM APIs + embeddings\nSQL / NoSQL databases\nCloud infrastructure (AWS)\nData pipelines and streaming systems\nVector Databases\nFor applications and inquiries, contact: hirings@openkyber.com","datePosted":"2026-08-04T16:44:26.798Z","dateModified":"2026-08-04T16:44:26.798Z","hiringOrganization":{"@type":"Organization","name":"Openkyber","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Alaska","addressRegion":"MI","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"abd103cd6da29b1ab77e3003"},"url":"https://jobsearcher.com/jobs/abd103cd6da29b1ab77e3003"}}