Go back

Applied Machine Learning Engineer (Model Layer)

Location: Remote (Open to candidates in any country)
Industry: AI / Consumer Software / LLM Applications
Employment Type: Full-time / Long-term dedicated contractor
Reports To: Executive Leadership / Founders
Work Schedule: Flexible, with some overlap for team collaboration
Compensation: Up to USD $2,000 per month

About Our Client
Our client is building a simple, private chat product that gives people easy access to powerful open-source AI models—with no technical setup required. The product is designed to be fast, affordable, and high-quality, connecting users directly to the best available models for their needs. The website and front-end are handled by a separate team, so this role is purely focused on the AI engine room.

They are a small, early-stage team looking for a hands-on engineer who wants to own a critical layer of the product—not just contribute to it.

The Role
The model layer is the engine room behind the chat. It is the part that:

Connects to the AI models.

Decides which model should answer each message.

Sends the request, returns the reply, and does it all while keeping answers fast, affordable, and good.

You will own this layer from front to back. You will work directly with the application team to integrate the model layer into the product, but your focus is the AI—not the front end.

This is a hands-on, production-focused role for someone who has built with large language models before—not just in notebooks, but in real products that real people use.

What You'll Do
Model Integration & Curation

Integrate open models through provider APIs and keep the lineup current as better models appear.

Evaluate new models as they are released and determine whether they should be added to the routing pool.

Model Routing & Selection

Build the routing logic that sends each request to the best model based on:

Quality

Refusal behavior

Speed

Cost

Continuously refine routing decisions using real data and feedback.

Evaluation & Measurement

Stand up evaluation systems to measure model quality and routing choices from real usage data.

Define and track key metrics that determine success (e.g., answer quality, latency, cost per request).

Prompt & Parameter Tuning

Tune prompts, parameters, and fallbacks to get the best result for each kind of task.

Optimize model behavior for different use cases and user needs.

Cross-Team Collaboration

Work with the application team to wire the model layer into the product and the policy layer.

Ensure seamless integration between the model layer, the product, and the user experience.

Performance & Cost Optimization

Watch latency and cost per request—and bring both down over time.

Continuously improve the efficiency of the model layer without sacrificing quality.

What We're Looking For (Requirements)
Production LLM Experience: Real experience building with large language models in production—not just in notebooks or prototypes.

Strong Python: Comfort working directly against model provider APIs (e.g., OpenAI, Anthropic, Together, Groq, or open-weight model providers).

Model Selection & Evaluation Judgment: Good instincts on model selection, evaluation, and prompt design.

Data-Driven Decision Making: Able to measure what matters and make decisions based on the numbers.

Startup Mindset: Happy in a small team, owning a layer from front to back. You are self-directed, proactive, and comfortable with ambiguity.

Nice to Have (Preferred)
Model Routing Experience: Experience with model routing, ensembles, or multi-model systems.

Open-Weight Model Familiarity: Familiarity with open-weight model families such as Llama, Qwen, or Gemma.

Inference Optimization: Background in inference cost and latency work.

Evaluation Frameworks: Experience building or using evaluation frameworks for LLM outputs.

Privacy-First AI: Interest or experience in privacy-focused AI products.

The Ideal Candidate (Mindset & Attributes)
Builder, Not Researcher: You want to build and ship production systems, not publish papers. You care about what works in the real world.

Pragmatic: You make practical trade-offs between quality, speed, and cost. You know when "good enough" is the right call.

Ownership Mentality: You take full responsibility for the model layer—its performance, its reliability, and its cost-efficiency.

Data-Driven: You don't guess; you measure. You use real usage data to inform routing and tuning decisions.

Curious & Current: You stay up to date on the latest models and techniques, and you're always looking for ways to improve.

Collaborative: You work well with the application team and can communicate technical concepts clearly to non-ML stakeholders.

What Success Looks Like
Success in this role is defined by measurable improvements in the model layer's performance. You will know you are succeeding when:

Model Routing Is Smart: Requests are consistently routed to the best model for the task—balancing quality, speed, and cost.

Answer Quality Is High: Users are getting accurate, helpful, and relevant responses from the models.

Latency Is Low: Response times are fast and consistent, even as usage grows.

Cost Per Request Is Optimized: The model layer is cost-efficient, and costs are trending downward over time.

Evaluation Is Data-Driven: You have systems in place to measure model quality and routing decisions from real data—not guesswork.

The Model Lineup Stays Current: New and better models are integrated quickly, and outdated models are retired.

Integration Is Seamless: The model layer works flawlessly with the product and policy layers.

You Own It End-to-End: The model layer is reliable, well-documented, and maintained without constant firefighting.

Why This Role Stands Out
Own a Critical Layer: You will own the engine room of the product—the part that makes everything work.

Work with Cutting-Edge Models: Engage with the latest open-weight and provider models as they are released.

Small Team, Big Impact: Join an early-stage team where your work directly shapes the product and user experience.

Remote & Global: Work from anywhere in the world with a flexible schedule.

Privacy-First Mission: Be part of a product that prioritizes user privacy and accessibility.

Growth Opportunity: As the product scales, so will your role and responsibilities.

Applied Machine Learning Engineer (Model Layer)

Job Category

Engineering

Job Type

Full Time (35 hours or more per week)

Work Schedule and Timezone

Sydney, NSW

Published on

Sep 13 2026

BruntWork will never ask you for money or any other form of payment. If someone claiming to represent BruntWork is requesting a payment from you, please let us know at applications@bruntwork.co

“BruntWork made the entire recruitment process smooth, transparent, and stress-free. They matched me with a client that genuinely fits my skills and values — and the support didn’t stop at placement... A reliable, professional partner I’d recommend without hesitation.”

— Zyrrah D, Bookkeeper

Google rating
4.9/5
Glassdoor rating
4.9/5