Do you ever feel like your computer testing is slow? Do your machine learning tools not use your computer power in the best way? This is a big problem for many AI companies today. Big AI models need huge computer systems to train. The people who build these computer systems have some of the most important jobs in tech right now.
This article talks about a company called Thinking Machines Lab. It defines a duty called “Research Infrastructure Engineer.” We will seek at what this duty is, what skills you want, how much money you can make, and how to apply for this job.
What Does This Job Do?
At Thinking Machines Lab, there is a team called Research Infrastructure. This team helps make the whole research process faster and better. One of the tools this team builds is called Tinker. Tinker is like a helper program. It lets researchers change and improve Artificial Intelligence models without needing to control the actual computers by hand.
Here is a simple way to picture how the systems work together:
- At the top: Researchers and scientists do their work, like training models and building coding tools.
- In the middle: The Thinking Machines Infrastructure sits here. It handles the Tinker tool, testing, and organizing tasks across many computers.
- At the bottom: The actual computer hardware sits here. This includes many powerful graphics cards (GPUs) working together.
The main job of these engineers is to stop slowdowns. They organize many computers so they can work together well. They also work on making training smarter, especially for a method called Reinforcement Learning. This method has different types, like RLHF, PPO, and DPO. These are ways to teach AI models using rewards and feedback.
These engineers also write very detailed, low-level code. This code helps computers do math quickly using short data formats, like FP8, BF16, NVFP4, and INT8. Short cut data formats can build training quickly and use less memory.
Another big part of the job is keeping long training jobs safe. Sometimes training a big model can take some days. During that time, hardware can break or have issues. So engineers build systems that save progress often. If something breaks, the training does not have to start over from zero.
What Jobs Are Open?
Thinking Machines Lab mostly works from San Francisco, California. Some jobs also have options in New York City. Here are some of the jobs open in 2026:
- Research Engineer, Infrastructure, Training Systems — Works on spreading training across many computers and organizing computer clusters. Location: San Francisco (partly remote).
- Research Engineer, Infrastructure, Kernels — Works on writing quickly, low-level code for GPUs and TPUs. Location: San Francisco (partly remote).
- Research Engineer, Infrastructure, Numerics — Works on using smaller number formats to save memory and speed up work. Location: San Francisco (partly remote).
- ML Infrastructure Engineer — Works on keeping computer clusters reliable and organizing training across many machines. Location: San Francisco or New York City (in-person).
- Software Engineer, Evaluation Platform / Infra — Works on testing tools that check how well AI models perform. Location: San Francisco (in-person).
What Skills Do You Need?
To get one of these jobs, you should be strong in four main areas.
Working with Many Computers at Once
You should know how to use tools that help spread AI training across many computers. Examples include PyTorch, JAX, DeepSpeed, and Megatron-LM. You should also know how to use tools like Slurm, Kubernetes, or Ray. These tools help manage many computers working together. You should also know how to save progress during training and how to fix problems automatically when a computer fails.
Programming Languages and Computer Basics
You need to know Python, C++, and Go. These are computer languages used to build fast and reliable systems. You should also understand how computers manage memory and how they send information to each other over a network.
Working with Computer Hardware
You should understand how to use smaller number formats like FP8, BF16, NVFP4, and INT8. These formats help save memory and make the computer work faster. You should also know how to help computers “talk” to each other quickly using tools like NCCL, MPI, and RDMA over Converged Ethernet (also called RoCE).
Understanding the Type of Work
You should understand how to build fast systems that create practice data for AI models to learn from. You should also know how to build fair tests to check if an AI model is good at things like coding, math, and reasoning.
How Much Does This Job Pay?
This job pays very well. Here is what you can expect:
- Base pay: Between $350,000 and $475,000 per year, or more. The exact amount depends on your experience level.
- Company shares: You may also get shares in the company. Since the company is still growing, these shares could become worth a lot more later.
- Location: You can work in San Francisco or New York, and many roles let you work from home part of the time.
- Visa help: The company can help workers from other countries get a visa. This includes the H-1B visa, the O-1 visa (for people with special skills), and OPT for people who studied in the science and math fields (STEM).
The company also offers full health coverage, money to help you keep learning, extra money for building your own computer setup at home, and flexible schedules that let you work from home sometimes.
How to Apply for the Job
If you want to apply, here are three steps to follow.
Make your resume clear and focused on results. Instead of just listing your skills, show real numbers. For example, you could write: “I made computer communication 35% faster across 512 GPUs by writing custom code.” Numbers like this help show your work clearly.
Learn about the company’s tools and projects. Look at any free code the company has shared online. Learn about their tools, like Tinker. This shows that you understand what the company does and that you are truly interested.
Apply through the right websites. Send your application through the official Thinking Machines job page, which uses a system called Ashby. You can also look for open jobs listed on investor websites, like Andreessen Horowitz (a16z) and Accel.
Common Questions
-
What makes Thinking Machines Lab different from other AI companies?
The company tries to connect its research team and its engineering team very closely. This means infrastructure engineers get to work directly with researchers who are building new kinds of AI models.
-
Do I need a PhD to get this job?
No, you do not need one. While some research jobs may prefer higher degrees, infrastructure jobs care more about your skills in building systems, working with many computers, and writing fast, low-level code. Your real experience matters more than your school degree.
Disclaimer: This article is written only to share general information. It is not made by or officially connected to Thinking Machines Lab, Inc. The duty titles, pay amounts ($350,000–$475,000+), and requirements come from public information found online. Before applying, please check the real and current job listings on the official Thinking Machines Ashby careers page.