justbacked
← All jobs
eComID

Senior Machine Learning Infrastructure Engineer

Stockholm, Sweden

Seed · $17M · 43d ago Posted 28d ago

About eComID


eComID is building the Shopping Passport for modern commerce - a shopper context layer that enables brands to understand and personalize every shopper from their very first visit. By combining AI-powered sizing, conversational search and intelligent product discovery, we’re making online shopping smarter, more personal and less wasteful.

Launched in Stockholm in 2024, eComID is already operating at scale, reaching more than 20 million shoppers every month. We recently raised a $17.4 million seed round - the largest in European fashion-tech history and the second-largest globally.

We’re backed by H&M Group, Stadium, leading international VCs and some of Europe’s most successful industry leaders and technology founders, including Helena Helmersson, Sebastian Knutsson, Maria Raga and Alan Mamedi.

Now, we’re bringing together exceptional people with warm hearts to build a world-class fashion-tech company from Stockholm. If you want to solve meaningful problems, move fast and help shape how the world shops, this is the place to do it.

What you will work on

We are looking for a Senior MLOps Engineer to own and evolve the systems that take our models from development into production.

You will sit at the intersection of machine learning, software engineering, and infrastructure. Your job is to make it easy for our AI engineers to build, deploy, evaluate, monitor, and improve models while ensuring those systems remain reliable, scalable, observable, and efficient in production.

This is not a role focused only on maintaining infrastructure. You will work directly with the engineers building our models and products, helping shape both the platform and the way we develop AI systems at eComID.

You will help build and own the infrastructure and tooling behind our production AI systems, including:

  • Building and improving model training, evaluation, and deployment pipelines

  • Designing reliable production systems for real-time and batch model inference

  • Improving how models move from experimentation to production

  • Building tooling that makes it easier and safer for AI engineers to ship new models

  • Model versioning, reproducibility, rollbacks, and release strategies

  • Monitoring model performance, latency, errors, resource usage, and production health

  • Designing automated evaluation and validation before and after deployment

  • Scaling inference workloads while balancing latency, reliability, and infrastructure cost

  • Managing feature and model dependencies across training and inference

  • Improving CI/CD workflows for machine learning services and pipelines

  • Building observability across models, pipelines, services, and infrastructure

  • Designing systems for experimentation, shadow deployments, A/B testing, and gradual rollouts

  • Making our AI platform simpler, faster, and more reliable as both our team and product surface grow

You will work closely with AI engineers to understand where the development process is slow, fragile, or repetitive and turn those problems into better infrastructure and tooling.

What we are looking for

We are looking for someone with strong software and infrastructure fundamentals who has experience operating machine learning systems in production.

You should have experience with several of the following:

  • Building and operating production ML infrastructure

  • Deploying and serving machine learning models at scale

  • Designing CI/CD systems for ML or backend services

  • Cloud infrastructure, ideally GCP

  • Containerised workloads and orchestration

  • Python in production environments

  • Infrastructure as code and automated environment management

  • Monitoring, observability, alerting, and production debugging

  • Model versioning, experiment tracking, and reproducibility

  • Training and inference pipelines

  • Distributed workloads and asynchronous processing

  • Real-time, low-latency inference systems

  • GPU workloads and inference optimisation

  • Reliability engineering and production incident handling

  • Designing internal platforms or developer tooling

Experience working directly with ML engineers or data scientists is important. You should understand enough about the ML lifecycle to identify where infrastructure can improve experimentation, deployment, evaluation, and production reliability.

We care more about strong engineering judgement and experience operating real systems than whether you have used our exact stack before.

Our current tech stack

Today our core stack includes:

Languages: Python, Go, TypeScript

Services/libraries: GCP, BigQuery, ClickHouse, PostgreSQL, Redis, Apache Beam, Kubeflow, Feast, PyTorch, Typesense, Milvus, ConnectRPC

We are not attached to technologies for their own sake and are open to introducing better tools where they make sense.

What we value

We are looking for someone who:

  • Takes ownership of the systems they build

  • Cares about reliability and understands what it means to operate systems in production

  • Enjoys working directly with AI engineers and product teams

  • Thinks carefully about failure modes, observability, reproducibility, and operational complexity

  • Can reason about trade-offs between latency, throughput, reliability, cost, and developer experience

  • Will question existing architecture and propose better approaches

  • Prefers simple, robust systems over unnecessary complexity

  • Automates repetitive work rather than accepting it as part of the process

  • Treats infrastructure as a product used by other engineers

  • Cares about making the path from experiment to production fast and predictable

  • Is comfortable debugging problems across application code, infrastructure, networking, data, and models

  • Thinks about the full lifecycle of a model, not only how to deploy it

  • Can build abstractions without hiding the underlying system when things go wrong

  • Knows when to introduce platform capabilities and when a simpler solution is enough

What we offer

  • A best-in-class team with warm hearts and no ego. We care deeply about the quality of our work, while staying kind, humble, and easy to work with. The goal is always to build the best product possible, not to win internal arguments.

  • A direct connection to product. You will work with product every day, understand why we are building something, help shape the solution, and own the delivery from idea to production.

  • Real ownership. You will have meaningful influence over our ML platform, infrastructure, tooling, and engineering decisions. We expect senior engineers to improve the way we work, not just maintain what already exists.

  • Work directly with the AI team. You will sit close to the engineers building and experimenting with our models, helping turn research and prototypes into reliable production systems.

  • Build infrastructure that directly changes how fast we can innovate. Improvements you make to deployment, evaluation, observability, and tooling will directly affect how quickly the team can experiment and ship.

  • Interesting technical problems with real users behind them. You will work on production AI systems where latency, reliability, scalability, and model quality all matter.

  • Freedom to choose the right tools. We have a strong existing stack, but no attachment to technology for its own sake. If there is a better way to solve a problem, we want to hear it.

  • High standards without unnecessary process. We care about reliability, thoughtful engineering, and maintainable systems while keeping teams small, communication direct, and decision-making fast.

  • The chance to shape the next stage of the company. We are still early enough that the platform, principles, and engineering practices you establish now will have a lasting impact on how eComID builds and operates AI.

Benefits

  • Health allowance: 5,000 SEK wellness allowance per year

  • Office: Work from our office in central Stockholm

  • Office perks: Weekly breakfasts and snacks throughout the day

  • Tech: Your choice of computer, phone, and phone contract

  • Learning & growth: Annual budget for courses, books, and conferences

  • Insurance: Extensive coverage, including private healthcare insurance

  • Pension: Private pension insurance via SPP