Senior ML Research Scientist, Perception Models
TwelveLabs· Remote·
Who we are
Video is 90% of the world's data. Most of it is invisible to machines.
TwelveLabs builds the intelligence layer to change that. Our multimodal AI models understand video the way humans do — across sight, sound, and motion — and power production-scale AI workloads across media, entertainment, sports, security, and government.
We have raised more than $210 million from NEA, Radical Ventures, Amazon, NVIDIA, Snowflake, Databricks, Index Ventures, NAVER Ventures, Korea Investment Partners, Quadrille Capital, Red Bull Ventures, and AI pioneers including Fei-Fei Li, Silvio Savarese, and Alexandr Wang.
We are a global company, headquartered in San Francisco with offices in Seoul, New York, and London, and employees around the world. We believe the differences in our cultural, educational, and life experiences make our products stronger. Building technology that understands the world in all its complexity requires people who see it from every angle. We are looking for individuals who are driven by hard problems and want their work to matter. Come build it with us!
About the Team
TwelveLabs builds multimodal foundation models that understand what happens in video and how visual information is organized across time and space. The Perception Models team develops general-purpose representations that work across full videos, clips, regions, and individual entities.
Rather than building a collection of narrow, task-specific models, we use signals from capabilities such as detection, segmentation, and tracking to strengthen shared video representations.
About the Role
As a Senior ML Research Scientist, you will lead research at the intersection of spatiotemporal modeling and multimodal representation learning. You will define open-ended research problems, validate ideas through rigorous large-scale experiments, and turn promising results into production capabilities.
This role is ideal for a hands-on researcher with deep expertise in at least one of these areas and a strong interest in connecting them.
In this role, you will
Research multimodal representations that connect temporal and spatial structure with semantic understanding.
Develop models that operate across multiple levels of granularity, from full videos and clips to regions and entities.
Design training objectives, datasets, evaluation methods, and large-scale experiments.
Use task-specific vision models as supervision or components, and integrate useful signals into general-purpose representations.
Partner with research and engineering teams to scale and ship new capabilities.
Even if you don't check every box, we encourage you to apply.
If you're a zero-to-one achiever, a ferocious learner, and a kind team player who motivates others, you'll find a home at TwelveLabs.
You may be a good fit if you have
A strong research track record in computer vision, video understanding, multimodal learning, or a related field.
Deep expertise in at least one of the following: video foundation models, self-supervised or contrastive learning, embeddings, and retrieval; or temporal modeling, object-centric learning, detection, segmentation, and tracking.
Strong hands-on skills in Python and PyTorch or an equivalent deep learning framework.
Proven ability to design and run rigorous experiments at scale.
Research impact through publications, production systems, or both.
The ability to own ambiguous problems and collaborate across research and engineering.
We evaluate based on relevant technical skills and sustained industry impact. This role is typically a strong fit for engineers with an MS and deep industry experience who have evolved from individual contributor to technical leader in production ML environments.
Preferred Qualifications
Experience with open-vocabulary vision or dense and region-level representations.
Experience with large-scale video training, data curation, or evaluation infrastructure.
Experience translating research into production.
Experience leading research projects or mentoring researchers.
What makes this role unique
This is an opportunity to define how a foundation model understands the same video across multiple levels of granularity, not simply improve a single task-specific benchmark.
Others
Work Location: Seoul Itaewon office + Pangyo satellite office
Benefits and Perks
Growth & Tools
글로벌 B2B 고객과 함께 성장하는 Global Team
자율성과 협업을 모두 갖춘 하이브리드 근무
최신 맥북 및 70만 원 상당 재택근무 장비 지원, 3년 주기로 최신 장비 교체
Tokens never sleep - Tech 직군 LLM 토큰 무제한 지원
강의, 컨퍼런스, 멤버십 등에 사용 가능한 연 140만원 상당 자기개발비 지원
영어 교육 프로그램 및 글로벌 버디 프로그램 운영
야간 및 주말 출퇴근 택시비 지원
Meal & Snack
식비·교통비 등 자유롭게 사용할 수 있는 연 720만원 상당 법인카드 제공
사무실 내 스낵바 운영 (간식, 커피, 제철 과일 등)
사무실 근무 시, 오후 7시 이후 저녁 식대 제공
Wellness & Family
연 1회 본인 및 가족 1인의 건강검진 제공
단체보험 가입 (상해보험/치아보험/가족 상해보험 중 택 1)
독감 예방접종비 지원
연말 2주간 유급 Holiday Break 운영