TixelJobs

Interview questions

Data scientist interview questions (2026)

Data scientist interviews tend to cover statistics and experimentation, SQL and analysis on a realistic dataset, a modeling or case discussion, and a presentation of past work to a mixed audience. The strongest candidates show judgment about what to measure and why, not just how to fit a model. TixelJobs lists 1,414 data scientist roles as of September 2026.

Updated September 22, 2026

About you and your background

  1. 01Tell me about your background and the kind of data science work you have done most.

    What they are checking: The interviewer wants to place you on the spectrum from analytics to modeling and see whether it matches the role.

    Example answer

    I have spent five years in data science, mostly in product analytics and experimentation at consumer companies. My first role was on a growth team, where I built the metrics layer and ran most of the onboarding experiments. Later I moved to a marketplace, where the work shifted toward causal questions: whether a pricing change caused the demand drop or whether it was seasonal, whether a new seller tool actually improved fulfillment or just attracted better sellers. I have built predictive models when a decision needed one, including a churn model and a lead scoring model, but the majority of my impact has come from getting the measurement right and making the tradeoffs legible to the people deciding. I write SQL and Python every day and I present to product and leadership most weeks.

  2. 02Describe an analysis that changed a decision. What was the decision and what did you find?

    What they are checking: They are checking whether your work actually influences outcomes or just produces reports.

    Example answer

    A product team wanted to expand a free trial from seven to fourteen days because a competitor had. I was asked to size the upside. When I looked at the historical conversion data, almost all trial conversions happened in the first four days or in the last two before expiry. Extending the trial would mostly delay the second cluster by a week and reduce cash collected in the quarter. I built a simple model of conversion timing under both trial lengths and showed the expected revenue delay alongside the best-case conversion lift, which was small. The team decided to keep seven days and instead test a reminder sequence at day five, which raised conversions in the experiment. The analysis mattered because it reframed the question from what the competitor did to when our users actually decide.

  3. 03What tools and languages do you use daily, and which have you dropped?

    What they are checking: A practical check on your stack plus a signal about whether you choose tools for the team or for yourself.

    Example answer

    Daily, it is SQL against the warehouse, Python with pandas and statsmodels for analysis, and a notebook environment for exploration that I move into versioned scripts when something becomes recurring. I use a BI tool for dashboards that other people need to self-serve, and I have written dbt models when the metrics layer needed a definition that would survive me. I dropped R a few years ago, not because it is worse but because the teams I worked on standardized on Python and being able to hand code to an engineer mattered more than my preference. I also stopped building dashboards for questions that will only be asked once. A short written memo with the numbers and the recommendation gets read; a dashboard nobody asked for gets ignored.

  4. 04How do you explain a complex result to someone without a technical background?

    What they are checking: Communication is half the job, and they want to hear a concrete method rather than a claim that you are good at it.

    Example answer

    I start from the decision they need to make and work backward to the minimum they need to understand. For a churn model I did not explain the algorithm; I showed which three behaviors were the strongest early warnings and what the model got wrong, using a handful of real accounts they knew. I use one chart per point, plain axes, and I say the number out loud before showing it so the chart confirms rather than surprises. I am explicit about uncertainty in words rather than intervals: this is likely, this we cannot tell yet. And I ask them to repeat back the conclusion, because if they cannot explain it to their own manager, I have not finished. The hardest part is leaving out things I found interesting that do not change the decision.

  5. 05What is a mistake you made in an analysis, and what did you change afterward?

    What they are checking: They want honesty about error and evidence that you built a habit to prevent it recurring.

    Example answer

    Early in my career I reported that a new landing page had lifted sign-ups by a meaningful amount. The test had run for five days, and I had not noticed that the variant had launched on a Monday and the comparison window included a weekend for the control. When the test ran a full two weeks the effect disappeared. I had to go back to the team and retract the recommendation, which was uncomfortable. Since then I decide the test duration and the analysis plan before launch, I never read results before the planned end unless there is a guardrail breach, and I check day-of-week balance as a routine step. The bigger change was cultural: I now describe results as provisional until the analysis plan is complete, and I write that plan down where others can see it.

Analysis and experimentation

  1. 06How would you design an A/B test for a change to the checkout flow, and how would you decide it worked?

    What they are checking: This tests the full experiment lifecycle: metric choice, randomization, power, duration and the decision rule.

    Example answer

    I would first define the primary metric, probably completed checkout rate per session that reached checkout, and a couple of guardrails such as refund rate and average order value. Randomization would be at the user level so a person sees a consistent flow, with assignment logged at the point they enter checkout, not at page load. I would run a power calculation with the baseline rate and the minimum effect worth shipping, which sets the sample size and therefore the duration, and I would run at least two full weeks to cover weekly cycles. Before launch I would run an A/A check on the assignment. At the end I would compare the primary metric with a confidence interval, check guardrails, look at a small number of pre-registered segments, and recommend shipping only if the interval clears the minimum effect.

  2. 07A key metric dropped 8% week over week. Walk me through how you investigate.

    What they are checking: They are looking for a disciplined investigation that rules out measurement problems before hunting for causes.

    Example answer

    First I check whether the drop is real: was there a tracking change, a pipeline delay, a definition change, or a partial day in the window. Many of these turn out to be measurement. If the number holds, I decompose it. I split by platform, region, new versus returning users, and acquisition channel to see whether the drop is broad or concentrated, and I check the funnel steps upstream to find where the leak starts. Then I look at what changed in that window: releases, experiments, marketing spend, pricing, outages, and external events like holidays. I write up what I have ruled out as I go so the team is not chasing the same theories. Once I have a leading hypothesis I look for evidence that would disprove it before I name it as the cause.

  3. 08When is a p-value the wrong thing to report, and what would you use instead?

    What they are checking: A statistics check that also reveals whether you understand what stakeholders need from a result.

    Example answer

    A p-value tells you how surprising the data would be if there were no effect. It does not tell you the size of the effect or whether it matters, so when the decision depends on magnitude, I report the effect estimate with a confidence interval and translate it into the unit the stakeholder cares about, like revenue per week. It is also the wrong thing when the sample is large enough that trivial differences are significant, when many metrics are being checked at once without correction, or when someone is peeking at a running test. For sequential decisions I would use a method designed for it, such as sequential testing or a Bayesian approach with a decision threshold. For exploratory analysis I would not report p-values at all; I would describe patterns and propose the experiment that would test them.

  4. 09How do you handle a situation where you cannot run a randomized experiment?

    What they are checking: This checks your causal inference toolkit and whether you state assumptions honestly.

    Example answer

    It happens often: pricing changes that roll out to everyone, a policy change in one region, a feature a partner will not let us withhold. My first choice is to find a natural comparison. A difference-in-differences design works when the change hit some units and not others and their trends were parallel before. A regression discontinuity works when eligibility is set by a threshold. When neither applies, I use matching or a synthetic control built from similar units that were not exposed. In every case I write down the assumptions the estimate depends on and try to falsify them, for example by checking for a fake effect in a period before the change. I report causal claims from observational data with more caution than experiment results, and I say what experiment would confirm them if the decision is important enough.

  5. 10How do you choose and validate a model when the goal is a decision rather than a prediction?

    What they are checking: They want to see that you evaluate models on the decision they support, with validation that matches how the model will be used.

    Example answer

    When the output drives a decision, I evaluate the decision, not just the prediction. For a lead scoring model, the question was which leads a small sales team should call first, so I measured revenue captured in the top decile rather than overall accuracy, and I compared against the rule the team already used. I chose a model the team could interrogate, in that case a regularized logistic regression with a few interaction terms, because a sales lead who does not trust the score will ignore it. Validation was time-based, training on earlier quarters and testing on the next, since a random split would leak. Before rollout we ran it alongside the old rule for a month and compared outcomes. The model was slightly less accurate than a boosted alternative but better at the actual decision.

Behavioral and teamwork

  1. 11Tell me about a time a stakeholder wanted a result you could not support.

    What they are checking: This tests integrity under pressure and whether you can offer a constructive alternative rather than a flat no.

    Example answer

    A marketing lead wanted to report that a campaign had driven a large share of new sign-ups. The attribution was last-touch and the campaign had run during a period when organic traffic also rose, so I could not support that number. I explained why last-touch overstated it, and rather than just saying no, I offered what I could support: a holdout comparison from the regions where the campaign had not run, which showed a smaller but real lift. I put both numbers in the deck with a clear note on what each one measured. The marketing lead was frustrated at first but used the holdout number, and it held up when finance asked questions. After that we agreed on an attribution approach in advance of campaigns so this did not repeat.

  2. 12Describe a project where your first approach was wrong.

    What they are checking: They want to see how quickly you notice you are on the wrong path and what you do about it.

    Example answer

    I was asked why customer support tickets had risen and I assumed it was a product quality problem, so I spent a week correlating tickets with recent releases. Nothing lined up. When I finally looked at the ticket text, a large share were questions about a new billing page that was working as designed but confusing. My first approach was wrong because I started from a hypothesis instead of the data. I switched to a bottom-up categorization of a sample of tickets, which took two days and pointed straight at the billing page. The product team rewrote the page copy and tickets fell within a week. I now begin every why question by reading a sample of the raw records before I touch aggregates, and I say my initial guess out loud so I notice when I am anchored on it.

  3. 13Tell me about working with engineers to get an analysis into a product or pipeline.

    What they are checking: This checks whether you can make your work operational and collaborate across the engineering boundary.

    Example answer

    I built a model that flagged orders likely to be delayed, and it was only useful if the operations team saw the flag in their tool. I worked with the engineer who owned that tool to define a small contract: my pipeline would write a score and a reason code to a table each hour, and their service would read it. I wrote the scoring job as a tested Python package rather than a notebook so their CI could run it, documented the schema, and agreed on what happened when the job failed, which was to show no flag rather than a stale one. We shipped in three weeks. The lesson for me was to design the handoff early: engineers are happy to integrate a model when the interface is clear and the failure mode is safe.

  4. 14Give an example of prioritizing between several requests when everyone said theirs was urgent.

    What they are checking: They want to see a repeatable method for prioritization and clear communication of the outcome.

    Example answer

    In one week I had a leadership request for a revenue forecast, a product manager who needed experiment results before a launch decision, and a sales team asking for a customer list. All were called urgent. I asked each one what decision depended on the work and when that decision would actually be made. The launch decision was the following day and could not move, so it went first. The forecast was for a board deck in two weeks, so I agreed on a scope and a draft date. The customer list was a repeat request, so I built a saved query the sales team could run themselves, which took an hour and removed the request permanently. I sent one message to all three with the order and the reasons. Nobody was thrilled but everyone knew where they stood.

  5. 15Tell me about a time you had to deliver a finding that nobody wanted to hear.

    What they are checking: This tests courage, preparation and whether you can deliver bad news in a way that leads to action.

    Example answer

    A product team had spent a quarter on a redesigned onboarding flow and the experiment showed no improvement in activation, with a slight decline in one segment. I was the one presenting. I prepared by making sure the analysis was airtight, checking the assignment, the sample size and the segment result for a multiple comparisons issue, so that the conversation could be about what to do rather than whether the numbers were right. In the meeting I led with what the experiment had taught us about which onboarding steps users actually skipped, which was useful, and then gave the result plainly. The team was disappointed but nobody argued with the data. They shipped a much smaller change based on the step analysis, and it moved activation. I would rather be the person who tells the truth early.

Your fit and the role

  1. 16Why do you want to work here rather than at a company with a larger data team?

    What they are checking: They are checking that you understand what a smaller team means day to day and want it for the right reasons.

    Example answer

    At a company with a large data team I would likely own a narrow slice, and I have done that; it is a good way to go deep on one thing. Right now I want the opposite: a role where I am close to the decisions, where I can see a question through from the raw data to the change in the product. Your team is small enough that the data scientist in the room is the one who set up the metric, ran the test and made the recommendation. That accountability is what I want. I am also drawn to the specific problem space, since I have worked on similar questions before and I know how much good measurement can change the outcome. I would expect to build some of the foundations myself, and I am comfortable with that.

  2. 17This role is closer to the product than to research. How does that suit you?

    What they are checking: The interviewer wants to confirm you will be satisfied by applied, fast-moving work rather than methodological depth alone.

    Example answer

    It suits me well. The work I enjoy most is the kind where a product manager has a real decision to make and the analysis changes what they do. I like being in the roadmap meeting rather than reading about it afterward. Product-facing work means being fast and clear more often than being exhaustive, and I have learned to give a good answer in a day and a rigorous one in a week, and to be explicit about which one I am giving. I still care about method, and I would push back if speed meant reporting something I did not believe. But the appeal of this role is exactly that the output is a product decision, and I would measure my own success by whether the team makes better decisions than it did before I joined.

  3. 18What would you do in your first month to understand our data?

    What they are checking: They want a concrete onboarding plan that shows how you learn a new data environment.

    Example answer

    In the first month I would read before I write. I would map the main tables and how the core metrics are defined, and I would find out who owns each definition, because the answer is often nobody. I would take the five most-used dashboards and reproduce the headline numbers from raw data to see where they diverge, which is the fastest way to find the hidden joins and filters. I would sit with the people who generate the data, such as the engineers who instrument events and the ops team whose actions become rows, and ask what they distrust. And I would pick one small question the team already has and answer it end to end. By the end I would have a short document of what is solid, what is shaky, and what I would fix first.

  4. 19How do you want to grow in the next two years?

    What they are checking: This reveals whether your growth goals line up with what the role and the team can offer.

    Example answer

    Two directions. First, I want to get better at causal inference on observational data, because so many of the important questions cannot be answered by an experiment and I want to be the person the team trusts on those. Second, I want to grow my influence on what gets measured, not just how. That means being in the room when a feature is scoped so that the instrumentation and the success metric exist before launch. I am not looking to manage people yet, but I would like to mentor a junior analyst and to own a domain rather than a queue of requests. This role looks like it has both: a real analytical challenge and a seat close to the product decisions, which is where influence on measurement comes from.

  5. 20What do you need from a team to do your best work?

    What they are checking: They want to know your working conditions and whether the team can realistically provide them.

    Example answer

    I need access to the raw data, not just curated views, and a clear owner for each metric so that I can resolve a definition dispute quickly. I need product managers who will tell me the decision they are trying to make rather than the chart they want. And I need a manager who protects time for the analysis to be done properly, and who will back me when a result is unwelcome. In return I try to be easy to work with: I write things down, I give an answer with a confidence level rather than stalling, and I tell people early when a request is going to take longer than they hoped. I do not need a large team or an elaborate stack; I need trust, access and a reasonable pace.

Questions worth asking them

  • How are experiments run today, and who decides when a result is good enough to ship?
  • Which metrics does the company steer by, and who owns their definitions?
  • What is the split between analysis, experimentation and modeling for this role?
  • Can you describe a recent decision that changed because of something the data team found?
  • How does the team handle a request that cannot be answered well with the data you have?

How to prepare

  • Practice SQL until window functions, joins across grains and date logic are automatic, because most loops include a live query.
  • Prepare one experiment story and one observational study story with the assumptions, the pitfalls and what you would do differently.
  • Rehearse a ten-minute presentation of a past analysis for a non-technical audience and time it.
  • Read the company's public metrics or product pages and come with two questions the data could answer.

Frequently asked questions

What statistics do I need for a data scientist interview?

Hypothesis testing and confidence intervals, power and sample size, regression with an understanding of confounding, and the basics of causal inference such as difference-in-differences. You should be able to explain when a p-value misleads and what you would do instead. Bayesian methods come up at some companies. Depth on a few of these beats shallow coverage of many.

How different is a data scientist interview from an ML engineer interview?

Data scientist loops weigh experimentation, statistics and communication more heavily and typically include a SQL or analysis round and a presentation. ML engineer loops lean toward coding, system design and deployment. Some companies use the titles interchangeably, so read the job description and ask the recruiter what the rounds are before you prepare.

Should I bring a portfolio or take-home project?

A concise portfolio helps, particularly if your work is not public. One or two write-ups that show the question, the approach, the result and what you would change are enough. If a company sends a take-home, treat it as a communication test as much as a technical one: a clear notebook with a short summary beats a long one with every chart.

Where can I find data scientist jobs?

TixelJobs lists 1,414 data scientist roles as of September 2026 on its data-scientist-jobs hub, gathered from company career pages and updated daily. Browsing is free. Membership is $9 a year or $4.90 every three months and unlocks the apply link and the full description on every job. Cancel any time from your billing page. Filter by remote, seniority and location.

Every data scientist job and its apply link, one payment a year.

Browsing is free. Membership is $9 a year or $4.90 every three months and unlocks the apply link and the full description on every job. Cancel any time from your billing page.