{"schemaVersion":"jobsearcher.job.v1","id":"9c48dc0176cd545734d2658d","url":"https://jobsearcher.com/jobs/9c48dc0176cd545734d2658d","canonicalUrl":"https://jobsearcher.com/jobs/9c48dc0176cd545734d2658d","title":"Senior Software Engineer","description":"Software Engineer – Technical Advisor (Contract)Contract (W2) | Remote – United States | $150/hour | 30–40 hrs/week## Contract Details| Contract type| W2 Contractor (no C2C, no visa sponsorship) || Length| 6 months, with possibility of extension || Pay| $150.00/hour| **Schedule** | 30–40 hours/week, mostly asynchronous; must be reachable during US business hours for periodic syncs. Your working hours are local to you — not restricted to Pacific time. || **Location** | Remote — must be based in and legally authorized to work in the United States without sponsorship |## Our Process1. **Recruiter screen** — a conversation about the role, your background, and fit2. **Video Call**3. **CodeSignal technical assessment** — a ~90-minute industry coding assessment (score of 500+ required)4. **Paid 2-week trial period** — two paid evaluations at $150/hour, testing real code-review and rationale-writing work5. **Interview** with the client6. **Start work**## About the RoleOur client is a leading AI safety and research company building reliable, beneficial, and interpretable AI systems. They're currently focused on evaluating how frontier AI coding models perform on real, production-grade software engineering work — and they're looking for a senior engineer to help lead that evaluation effort.As a **Software Engineer – Technical Advisor**, you won't be writing product code from scratch. You'll be the high-bar technical authority who determines whether AI-generated code is genuinely sound engineering or just plausible-looking output. That means reviewing model-generated pull requests against real production repositories, reading full agent session logs to understand what a model actually verified versus assumed, and building hard, container-based test problems that expose exactly where and why frontier coding models break down.*(Note: the end client's name is disclosed during the recruiter screening call, but is kept confidential in this posting.)*## What You'll Do- Audit model-generated pull requests against real production repositories, documenting every issue you find along with its severity and a detailed technical rationale- Evaluate full coding-agent sessions to analyze what the model investigated, verified, assumed, or skipped- Design and build container-based (Docker) technical benchmarks used to test AI models- Write clear, original technical rationale explaining *why* code fails — all written work must be self-authored, without AI text generators- Collaborate asynchronously with AI researchers to share findings and help refine evaluation criteria## What We're Looking For- 8+ years of production engineering experience preferred (exceptions considered for clearly exceptional profiles)- Background as a Senior, Staff, or Principal Software Engineer, Tech Lead, or Open-Source Maintainer- Experience in production codebases with a strict code review culture — startup, big tech, or open source backgrounds are all a fit- Polyglot adaptability: heavy Python and TypeScript experience is common, but you're comfortable dropping into an unfamiliar language on short notice- Demonstrated, hands-on comfort with Docker, git, and the command line to reproduce, isolate, and debug issues locally- Cross-layer fluency — comfortable working across backend, frontend, APIs, data, testing, or developer tooling- Strong written communication skills — you can articulate *why* code fails, not just how to fix it## This Role Isn't a Fit If You...- Are an Engineering Manager or Director who no longer writes or reviews code regularly- Have experience limited to low-code, no-code, or single-framework web MVPs without production code review exposure- Are a pure QA tester or non-technical prompt engineer- Don't have hands-on Docker and terminal/CLI experience- Are unwilling or unable to complete a coding assessment or write original technical rationale## How to ApplySend your resume to **Kamran Khan** at **kamran.khan@leadstackinc.com**, or apply directly and we'll follow up to schedule a screening call.","company":"LeadStack","rawCompany":"leadstack","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-09-04T08:24:32.566Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1251.00","title":"Computer Programmers","slug":"computer-programmers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Software Engineer","description":"Software Engineer – Technical Advisor (Contract)Contract (W2) | Remote – United States | $150/hour | 30–40 hrs/week## Contract Details| Contract type| W2 Contractor (no C2C, no visa sponsorship) || Length| 6 months, with possibility of extension || Pay| $150.00/hour| **Schedule** | 30–40 hours/week, mostly asynchronous; must be reachable during US business hours for periodic syncs. Your working hours are local to you — not restricted to Pacific time. || **Location** | Remote — must be based in and legally authorized to work in the United States without sponsorship |## Our Process1. **Recruiter screen** — a conversation about the role, your background, and fit2. **Video Call**3. **CodeSignal technical assessment** — a ~90-minute industry coding assessment (score of 500+ required)4. **Paid 2-week trial period** — two paid evaluations at $150/hour, testing real code-review and rationale-writing work5. **Interview** with the client6. **Start work**## About the RoleOur client is a leading AI safety and research company building reliable, beneficial, and interpretable AI systems. They're currently focused on evaluating how frontier AI coding models perform on real, production-grade software engineering work — and they're looking for a senior engineer to help lead that evaluation effort.As a **Software Engineer – Technical Advisor**, you won't be writing product code from scratch. You'll be the high-bar technical authority who determines whether AI-generated code is genuinely sound engineering or just plausible-looking output. That means reviewing model-generated pull requests against real production repositories, reading full agent session logs to understand what a model actually verified versus assumed, and building hard, container-based test problems that expose exactly where and why frontier coding models break down.*(Note: the end client's name is disclosed during the recruiter screening call, but is kept confidential in this posting.)*## What You'll Do- Audit model-generated pull requests against real production repositories, documenting every issue you find along with its severity and a detailed technical rationale- Evaluate full coding-agent sessions to analyze what the model investigated, verified, assumed, or skipped- Design and build container-based (Docker) technical benchmarks used to test AI models- Write clear, original technical rationale explaining *why* code fails — all written work must be self-authored, without AI text generators- Collaborate asynchronously with AI researchers to share findings and help refine evaluation criteria## What We're Looking For- 8+ years of production engineering experience preferred (exceptions considered for clearly exceptional profiles)- Background as a Senior, Staff, or Principal Software Engineer, Tech Lead, or Open-Source Maintainer- Experience in production codebases with a strict code review culture — startup, big tech, or open source backgrounds are all a fit- Polyglot adaptability: heavy Python and TypeScript experience is common, but you're comfortable dropping into an unfamiliar language on short notice- Demonstrated, hands-on comfort with Docker, git, and the command line to reproduce, isolate, and debug issues locally- Cross-layer fluency — comfortable working across backend, frontend, APIs, data, testing, or developer tooling- Strong written communication skills — you can articulate *why* code fails, not just how to fix it## This Role Isn't a Fit If You...- Are an Engineering Manager or Director who no longer writes or reviews code regularly- Have experience limited to low-code, no-code, or single-framework web MVPs without production code review exposure- Are a pure QA tester or non-technical prompt engineer- Don't have hands-on Docker and terminal/CLI experience- Are unwilling or unable to complete a coding assessment or write original technical rationale## How to ApplySend your resume to **Kamran Khan** at **kamran.khan@leadstackinc.com**, or apply directly and we'll follow up to schedule a screening call.","datePosted":"2026-09-04T08:24:32.566Z","dateModified":"2026-09-04T08:24:32.566Z","hiringOrganization":{"@type":"Organization","name":"LeadStack","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"9c48dc0176cd545734d2658d"},"url":"https://jobsearcher.com/jobs/9c48dc0176cd545734d2658d"}}