JOBSEARCHER

Software Engineer, Infrastructure

Flexcompute is a cutting-edge technology startup that specializes in ultra-fast simulation technology. Our products are utilized by companies in designing and optimizing technology products, with applications ranging from designing airplanes and cars to wind turbines and quantum computing chips. Our customer base includes both household names and startups in emerging industries. Our company was founded by world-renowned leaders in simulation technology from Stanford University and MIT. Backed by top VC firms, we are poised to disrupt the billion-dollar engineering simulation industry with our fast-growing trajectory.Tidy3D is a GPU-accelerated electromagnetic simulation product delivered as a cloud service. Behind the solver sits the platform that makes everything work: the Python client, the API, the job submission and scheduling layer, and the web application.We are hiring a Software Engineer for the Tidy3D infrastructure team to own that platform. You will build new capability, keep the existing system healthy, and run the release and deployment process.ResponsibilitiesDesign and build backend services for the Tidy3D platform.Develop and operate the control plane: the APIs, services, and data model behind task submission, job state, and result deliveryBuild scheduling and resource management for simulation jobs across heterogeneous GPU capacityHandle the operational concerns that come with a multi-tenant product: authentication, authorization, usage metering, and quota enforcementHelp with the deployment of our products to customers.Manage packaging and release to PyPI, and keep client and backend versions compatible across a long tail of installed versionsOwn the release pipeline end to end: versioning, CI/CD, staged rollout, and rollbackStandardize deployment patterns so the same product ships to our cloud, to customer-managed cloud accounts, and to on-premises installationsKeep production healthy.Instrument the platform and own its monitoring and alertingRespond to incidents and debug across boundaries, from a customer's Python traceback down to a stuck job on a GPU nodeManage cloud cost and capacity as usage growsRequirements2+ years building and operating production cloud services. New grads with exceptional background are encouraged to apply tooStrong Python. You have written backend services in it, not just scripts. Hands-on experience with a major cloud provider such as AWS, Docker, and KubernetesInfrastructure as code in production, ideally Terraform, with reusable modules rather than hand-managed environmentsLinux fluency and comfort operating in productionNice to haveGPU or HPC workloads, cluster schedulers such as Slurm, Ray, or Kueue, or high-throughput batch computeDesigning, packaging, and distributing a developer-facing Python library or SDKOn-premises, self-hosted, or air-gapped software deployment, and the enterprise requirements that come with it: SSO, network isolation, security reviewA background in physics, engineering, or scientific computing, or prior work on technical software for technical usersOpen source contributions to infrastructure or scientific computing projectsBenefitsCompetitive compensation with equity of a fast-growing startupMedical, dental, and vision health insurance401(k) ContributionGym allowanceFriendly, thoughtful, and intelligent coworkers