Python / Data Science Essentials Interview Questions
1. What is NumPy and why is it significantly faster than plain Python lists for numerical work?
NumPy (Numerical Python) is the foundational library for scientific computing in Python. At its core it provides the ndarray  an N-dimensional array of a single, fixed data type stored in a contiguous block of memory....
2. What are the main ways to create NumPy arrays?
Knowing the idiomatic array-creation functions is a baseline NumPy skill. Each function is designed for a specific situation and picking the right one keeps code readable and avoids unnecessary copies....
3. How do NumPy array shape, reshape, and axis work?
Every NumPy array has a shape attribute — a tuple giving the size along each dimension. Shape is fundamental because most NumPy operations depend on it, and shape mismatches are the most common source of errors in numerical code....
4. What is NumPy broadcasting and how does it work?
Broadcasting is the set of rules NumPy uses to perform element-wise operations on arrays of different but compatible shapes, without physically copying data to make them the same size. It is one of the most powerful and often misunderstood NumPy...
5. How does NumPy boolean masking and fancy indexing work?
Beyond basic integer indexing, NumPy supports two advanced selection mechanisms that are essential for data-cleaning and filtering tasks. Boolean masking : A comparison on an array produces a boolean array of the same shape....
6. What are the most commonly used NumPy mathematical functions in data science?
NumPy ships a comprehensive set of universal functions (ufuncs) — compiled, vectorised operations that apply element-wise across the full array without Python loops. Knowing these avoids writing slow manual loops for standard computations....
7. What is a Pandas DataFrame and how does it differ from a NumPy array?
A Pandas DataFrame is a two-dimensional, labelled data structure — think of it as a spreadsheet or a SQL table in memory. Rows and columns both have labels (the index and the column names ), and each column can hold...
8. How do you read CSV, Excel, and JSON files into a Pandas DataFrame?
Pandas has a family of pd.read_* functions that handle virtually every common data format. Getting data in is usually the first step of any data science workflow, so these functions deserve close attention....
9. What is the difference between df.loc[] and df.iloc[] in Pandas?
This distinction is tested in almost every Pandas interview. The short version: loc selects by label ; iloc selects by integer position ....
10. How do you detect, handle, and fill missing values in a Pandas DataFrame?
Missing values are represented in Pandas as NaN (float Not-a-Number from NumPy), NaT (Not-a-Time for datetime columns), or pd.NA (the newer nullable integer/string missing marker). Handling them correctly is the most time-consuming step of real-world data cleaning....
11. What are the different ways to filter rows in a Pandas DataFrame?
Row filtering is one of the most frequent DataFrame operations. Pandas provides several syntaxes, each with different readability and performance trade-offs....
12. How does Pandas groupby work and what aggregation patterns are most useful?
GroupBy is the Pandas implementation of the split-apply-combine pattern: split the DataFrame into groups by one or more column values, apply an aggregation or transformation to each group, and combine the results into a new DataFrame. It is the primary...
13. How do you merge and join DataFrames in Pandas, and what do the different join types mean?
Real-world data lives in multiple tables. Pandas merge() implements SQL-style joins, and concat() stacks DataFrames....
14. When should you use df.apply() versus vectorised Pandas operations?
apply() runs a Python function on every row or column of a DataFrame. It is the most flexible transformation tool in Pandas but also the slowest because it falls back to a Python-level loop under the hood....
15. How do you use pd.pivot_table to summarise data?
pd.pivot_table reshapes and aggregates a DataFrame simultaneously, producing a cross-tabulation — exactly like a spreadsheet pivot table. It is the go-to function for producing summary reports broken down by two categorical dimensions....
16. How do you perform string operations on Pandas DataFrame columns?
Pandas exposes string methods through the .str accessor on object-dtype Series. These operations are vectorised over the whole column — no explicit loop needed — and handle NaN values gracefully (they propagate as NaN rather than raising an error)....
17. How do you work with dates and times in Pandas?
Time-series data is everywhere in data science — sales by day, sensor readings by second, user activity by hour. Pandas has first-class datetime support built on NumPy's datetime64 type and Python's datetime module....
18. What is Matplotlib and what are the key components of a figure?
Matplotlib is Python's foundational plotting library, originally modelled after MATLAB's plotting API. Almost every other Python visualisation library (Seaborn, Pandas .plot(), Plotly static exports) either wraps Matplotlib or uses it as a rendering backend.
19. What are the most common chart types in Matplotlib and when do you use each?
Choosing the right chart type communicates data clearly; choosing the wrong one obscures it. Here are the workhorses of exploratory data analysis: import matplotlib.pyplot as plt import numpy as np fig, axes = plt ....
20. How do you create multi-panel figures with Matplotlib subplots?
Multi-panel figures are standard in data science reports — comparing multiple variables or time periods side by side. Matplotlib provides several ways to arrange subplots....
21. What is Seaborn and how does it differ from Matplotlib?
Seaborn is a high-level statistical visualisation library built on top of Matplotlib. Where Matplotlib gives you full control over every pixel, Seaborn provides opinionated, attractive defaults and plot types designed specifically for statistical exploration — with far less boilerplate code....
22. What are the most important Seaborn plot types for exploratory data analysis?
Seaborn divides its plots into relational (relationship between variables), distributional (distribution of a single variable), and categorical (comparison across categories). Knowing when to use each makes EDA far more efficient....
23. How do you create and interpret a correlation heatmap with Seaborn?
A correlation heatmap is one of the first plots every data scientist makes on a new dataset. It shows the Pearson (or other) correlation coefficient between every pair of numeric features as a colour-coded grid, immediately revealing which variables move...
24. What is Seaborn's FacetGrid and how does it enable multi-panel statistical plots?
FacetGrid is Seaborn's mechanism for trellis/small-multiples plots — the same chart repeated across different subsets of the data, defined by one or more categorical columns. It is one of Seaborn's most powerful features for exploring interaction effects between variables....
25. How do you compute descriptive statistics on a Pandas DataFrame?
Descriptive statistics summarise the central tendency, spread, and shape of a dataset. Pandas df.describe() is the starting point for any exploratory analysis, but knowing the individual methods gives you more precise control....
26. How do you reduce a Pandas DataFrame's memory usage through dtype optimisation?
DataFrames loaded from CSV often use unnecessarily large dtypes — 64-bit integers for values that fit in 8 bits, generic object dtype for repeated string categories. Downcasting dtypes can reduce memory by 4–8× without any data loss, enabling analysis of...
27. How do you generate reproducible random data with NumPy?
Reproducibility is a core requirement of data science — experiments, train/test splits, and simulations must produce the same result every run so that results can be verified and shared. NumPy's random number generation is the building block for all of...
28. How do you use value_counts() and pd.crosstab() to understand categorical data?
Categorical columns are understood by counting their frequencies and cross-tabulating them against other variables. These two tools answer the questions 'what values exist and how often?'...
29. How do you style Matplotlib figures and save them for reports?
The default Matplotlib style is functional but plain. For presentations and reports you need publication-quality output — chosen colour palettes, correct font sizes, no chart junk, and lossless or high-resolution raster output....
30. What is np.where and how is it used for conditional array creation?
np.where is NumPy's vectorised if/else for arrays. In its three-argument form it returns a new array built element-by-element: where the condition is True, use values from x ; where False, use values from y ....
31. What is Pandas method chaining and how does df.pipe() support it?
Method chaining is the style of writing data transformations as a single expression where each step's result is the input to the next. It avoids creating intermediate variables, reads like a pipeline, and makes the data flow explicit from top...
32. What does a typical exploratory data analysis (EDA) workflow look like in Python?
EDA is the first thing you do with a new dataset before any modelling. The goal is to understand the data's structure, quality, and relationships, and to spot problems (wrong dtypes, missing values, outliers, data leakage) before they propagate into...
33. How do you stack, concatenate, and split NumPy arrays?
Combining and splitting arrays is a frequent operation in data preprocessing — assembling feature matrices from multiple sources, or splitting a dataset into folds for cross-validation. import numpy as np a = np ....
34. How do you detect and remove duplicate rows in a Pandas DataFrame?
Duplicate rows silently inflate counts, distort means, and can cause data leakage between training and test sets. Pandas provides duplicated() and drop_duplicates() for systematic duplicate management....
35. How do you control colours and colour palettes in Matplotlib and Seaborn?
Colour is one of the most impactful design decisions in a chart. Used correctly it encodes information; used poorly it confuses or misleads....
36. How do rolling and expanding window functions work in Pandas?
Window functions compute statistics over a sliding or expanding subset of rows, essential for time-series smoothing, trend detection, and feature engineering. Unlike groupby aggregations, window functions return a result for every row, preserving the original index....
37. How do Seaborn jointplot and pairplot help explore multivariate relationships?
When you have more than one numeric variable, the next step after individual histograms is to understand relationships between pairs. Seaborn's jointplot and pairplot automate this exploration with minimal code....
38. What are the key performance tips when using NumPy for large-scale data processing?
NumPy is fast by default, but a few common mistakes can undermine that speed. Knowing these patterns makes the difference between code that runs in seconds and code that runs in minutes....
39. How do you visualise regression results and residuals using Seaborn and Matplotlib?
After fitting any regression model, visualising the residuals (actual - predicted values) is mandatory. Patterns in residuals reveal model assumptions violations: non-linearity, heteroscedasticity, or non-normality of errors....
40. How do you process large CSV files that don't fit in memory using Pandas?
When a CSV is larger than available RAM, loading it with a plain pd.read_csv causes a MemoryError . Pandas provides three strategies: chunking, selective loading, and dtype optimisation....
41. How do you add annotations and text to Matplotlib charts?
Annotations turn a chart into a story — highlighting a key data point, marking a threshold, or labelling significant events on a timeline. Matplotlib provides ax.annotate() for arrow-and-text annotations and ax.text() for free-form text placement....
42. How do you quickly extract top/bottom rows and random samples from a Pandas DataFrame?
During EDA you often need to inspect extremes (the highest-revenue customers, the worst-performing products) or draw a random sample for quick analysis. Pandas provides concise methods for each of these....
43. How is NumPy linear algebra used in data science applications?
Linear algebra underpins almost all of machine learning — from computing gradients to PCA to solving systems of equations. NumPy's linalg submodule provides production-grade implementations of the core operations....
44. How do you compare distributions across categories using Seaborn categorical plots?
Comparing how a numeric variable's distribution differs across groups is one of the most common analytical tasks. Seaborn's categorical plot family gives you progressively more information from left to right: bar (mean only) → box (five-number summary) → violin (full...
45. How do you build an end-to-end data cleaning and visualisation pipeline with NumPy, Pandas, and Seaborn?
Combining all three libraries in a coherent pipeline is what data science interviews and take-home assignments test. Below is a realistic miniature pipeline that demonstrates the key integration points....