E
Emergentlabsinc29d ago
New
New
Data Scientist
Data ScientistData
3 views0 saves0 applied
Quick Summary
Key Responsibilities
Turn agent trajectories, support tickets, logs, and user prompts into structured,
Technical Tools
Data ScientistData
Responsibilities
~2 min read- →Turn agent trajectories, support tickets, logs, and user prompts into structured, queryable signal through summarize-then-embed-then-cluster pipelines (à la Anthropic's Clio and Braintrust Topics): distill each trace along a dimension with an LLM, embed the summary, cluster and name the patterns, then classify at scale
- →Surface early indicators of confusion, a coming bug wave, churn risk, or fraud that no dashboard would ever surface on its own, and route them to the right team
- →Build predictive models that forecast conversion, retention, expansion, and churn, and embed those signals directly into product and growth workflows
- →Own marketing attribution and MMM: build the media-mix and incrementality models that tell us what's actually driving signups and paid conversions when per-user attribution is partial and, on mobile, broken by design
- →Own product and growth analytics across the self-serve funnel, web and mobile, covering activation, engagement, retention, and conversion, and design and analyze A/B and growth tests with real rigor around power, novelty effects, interference, and causal inference
- →Run clustering pipelines over hundreds of thousands of agent trajectories to discover the recurring kinds of things users try to build and the recurring ways builds fail, then hand product a taxonomy nobody had to hand-label, along with which clusters predict churn
- →Model the "aha moment" for new users, including text-derived features from their first prompts and first agent interactions, and rebuild onboarding around the earliest signals of long-term retention
- →Build a gross-margin model that attributes LLM and compute cost down to the individual app and cohort, and tell product which segments are net-positive
- →Untangle a fraud ring that looks anomalous on compute spend but has real payment history, decide whether it's an enforcement problem or a pricing problem, and defend the call with the data
- 2 to 5 years in data science or applied ML with a focus on product analytics, growth, or user behavior
- Strong SQL and real comfort working with large, event-level behavioral data at scale
- Solid classical ML foundations, including clustering (k-means, HDBSCAN, hierarchical), embeddings and vector similarity, dimensionality reduction (UMAP/PCA), and classification, with a working understanding of when each is and isn't the right tool
- Genuine skill at deriving insight from unstructured natural-language data (LLM traces, logs, tickets, free text) and turning it into predictive, queryable signal. Familiarity with topic-modeling and trace-clustering approaches (Clio-style summarize-then-embed pipelines, BERTopic, c-TF-IDF) is a strong plus
- Proficient in Python and the standard data-science stack (pandas, scikit-learn, statsmodels, numpy)
- Data engineering competence: you can design and ship ETL and data models (dbt or equivalent), not just query what already exists
- Experienced designing and analyzing experiments: sample sizing, power, significance, novelty effects, interference between tests, and causal methods
- You move fast and go deep, turning around in hours the analysis that takes most people days, because you've built the intuition to get to the right answer and the discipline to pressure-test it before anyone else sees it
- You dig past the top-line number to find the confound, you ask whether the metric measures what everyone assumes it measures, and you never hand over a figure without saying what it supports, what it doesn't, and what you'd check next
- You use AI agents aggressively to multiply your output, but you treat every AI-assisted result as a draft, not a deliverable
- Deep curiosity about user behavior and a real instinct for what drives growth, retention, and abuse
- Able to move fluidly between exploratory analysis, ML modeling, hypothesis testing, and crisp strategic recommendations, and to translate all of it into narratives that drive decisions
Nice to Have
~1 min read- Marketing Mix Modeling (MMM), media attribution, or incrementality and geo-testing experience, especially in low-tracking or post-cookie environments
- Experience at a PLG company with a self-serve funnel and freemium or usage-based/credit-based pricing
- Modern data stack (BigQuery, dbt) and product analytics platforms (PostHog, Amplitude, Mixpanel, Segment)
- Causal inference methods (difference-in-differences, synthetic control, propensity score matching)
- Fraud, trust and safety, or abuse analytics
- Working knowledge of embedding models and vector search, and the practical tradeoffs of running them at scale
- Familiarity with the economics of AI/LLM products, including COGS modeling where compute is the dominant variable cost
- You've built or contributed to AI-powered analytical tooling or novel measurement approaches
- Senior and Staff data scientists who want one of the hardest, most consequential analytics surfaces in AI software
- Applied ML practitioners who'd rather build the pipeline than wait for one
- Analysts who moved into ML and never stopped shipping
- Anyone drawn to where classical ML, LLM-native insight extraction, margin, attribution, and product all collide
What We Offer
~1 min read✓Daily Meals: Lunch and Dinner provided
✓Family Insurance: 5 Lakhs worth of coverage for you and your family
✓Unlimited Paid Time Off: Take the time you need to recharge and come back refreshed
✓Flexible Working Hours: Work arrangements that fit your life and commitments
Location & Eligibility
Where is the job
Bangalore, India
On-site at the office
Who can apply
IN
Listing Details
- Posted
- August 18, 2026
- First seen
- August 18, 2026
- Last seen
- September 15, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 60%
- Scored at
- August 18, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application · ~5 min on Emergentlabsinc's site
Please let Emergentlabsinc know you found this job on Jobera.
Similar Data Scientist jobs
View all →Research Scientist - Information Theory and Statistical Inference
I
Ifm UsResearch Scientist – World Modeling, Data
USD 150000–400000
Full-time
Senior Data Scientist, Simulation Capacity Optimization
USD 213000-263000
Generative AI Engineer
Full Time - PermanentRemote
Data Scientist
G
GlsllcData Scientist - December 2026 - May 2027 Grads
Full-Time
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.
E
Data Scientist