Senior AI / Software Engineer (Python, OCR, LLM & GPU Systems)
Job DescriptionEducation: If you did your undergrad in India, save your time and DO NOT APPLYYou will have to submit a W-9 when hired, no OPT, W-8, or other types of sponsorshipSenior AI / Software Engineer (Python, OCR, LLM & GPU Systems)NextGen Coding CompanyLocation: New York City — Hybrid / In-Person RequiredCompensation: $20–$25/hourEngagement Type: Part-Time, Approximately 20 Hours/WeekOpportunity: Potential to expand hours based on performance and project needsWork Authorization: Candidates must be legally authorized to work in the United States under an arrangement compatible with a W-9 contractor engagement. NextGen Coding Company cannot provide OPT or employment/visa sponsorship for the role.Role OverviewNextGen Coding Company is hiring a highly technical AI / Software Engineer in New York City to help build a large-scale document intelligence and AI platform.The engineer will work heavily with Python, OCR, document processing, data pipelines, LLMs, search/RAG, AWS, and NVIDIA GPU infrastructure.A critical requirement is supporting two versions of the platform:On-prem / air-gapped: running local models such as Qwen on NVIDIA H100/H200 infrastructureCloud: running in AWS with Claude, Qwen, OpenAI, Gemini, and other LLM/model integrationsThe underlying platform must be portable across both environments.ResponsibilitiesBuild production Python services, APIs, and data-processing pipelinesProcess large volumes of PDFs, HTML, images, scans, and structured/unstructured dataBuild OCR, extraction, parsing, normalization, and document-classification pipelinesExtract tables, entities, relationships, metadata, citations, and structured recordsBuild ETL and high-volume batch-processing systemsDevelop hybrid search, embeddings, vector search, RAG, reranking, and evidence retrievalIntegrate Claude, Qwen, OpenAI, Gemini, and open-source modelsDeploy and serve local LLMs on NVIDIA H100/H200 GPUsWork with vLLM, PyTorch, Hugging Face, CUDA, quantization, batching, and GPU inference optimizationBuild and maintain the AWS/cloud deploymentWork with PostgreSQL, OpenSearch, S3, Redis, queues, and related data infrastructureBuild APIs, webhooks, Stripe integrations, authentication, and third-party integrationsWrite automated tests covering OCR, extraction, retrieval, data pipelines, and AI outputsDebug complex AI, data, backend, and infrastructure issuesDocument architecture and implementation decisionsRequired ExperienceStrong Python engineering experienceStrong backend and data-engineering fundamentalsExperience building production softwareOCR / document intelligence experienceExperience processing PDFs, images, and large datasetsProduction experience with LLMs, RAG, embeddings, and vector searchExperience deploying open-source models on NVIDIA GPU infrastructureUnderstanding of H100/H200-class inference environmentsStrong AWS experiencePostgreSQL / SQLDocker and LinuxREST APIs and third-party integrationsAbility to independently own difficult engineering problemsTechnical StackCore: Python, FastAPI, SQL, PostgreSQL, RedisAI / LLM: Qwen, Claude, OpenAI, Gemini, Hugging Face, PyTorch, vLLMOCR / Documents: PaddleOCR, Tesseract, OpenCV, PyMuPDF, Docling or similarSearch / Data: OpenSearch/Elasticsearch, pgvector/vector databases, S3/MinIO, ETL pipelinesGPU: NVIDIA H100/H200, CUDA, vLLM, model quantization, local inferenceCloud: AWS, Bedrock, EC2, EKS/ECS, S3, RDS, OpenSearch, SQS, IAMInfrastructure: Docker, Kubernetes, CI/CD, LinuxProduct: REST APIs, webhooks, authentication, Stripe, third-party integrationsReact, Next.js, and TypeScript experience is helpful but not the primary focus.Ideal CandidateWe are looking for someone who can be given thousands of pages of PDFs, scans, HTML, images, and other data plus access to an H100/H200 server and an AWS environment — and independently help engineer the system required to process, OCR, structure, search, analyze, and query the information using modern AI models.The ideal candidate is stronger in Python, AI, data processing, OCR, backend systems, and infrastructure than traditional frontend development.You should be comfortable moving between an OCR pipeline, Python data processing, Claude/AWS integration, Qwen running on an H100/H200, databases, APIs, and production infrastructure.About NextGen Coding CompanyNextGen Coding Company is a U.S.-based software engineering firm building custom software, AI systems, automation platforms, and enterprise applications.Projects span AI/ML, document intelligence, data engineering, financial and compliance technology, cloud infrastructure, and enterprise software.Application RequirementsApplicants should provide:Resume or LinkedInGitHub and/or technical portfolioExamples of production AI/data systems builtExperience with Python, OCR, LLMs, AWS, and GPU infrastructureSpecific NVIDIA GPU/model deployment experience, if applicableQualified candidates will participate in a technical discussion and engineering review before engagement.