All roadmaps

Data Analyst

From raw spreadsheets to decisions a business will actually make.

The complete analyst path for 2026: what analytics actually is, spreadsheets, SQL, Python, statistics you can defend, dashboards, and AI-assisted analysis you can verify - every topic explained, with free resources to go deeper. Fresher and experienced tracks.

Your progress0/108 topics · 0%

Progress is saved in this browser.

11 stages108 topics₹0 to follow

Filled nodes are the focus for this track; outlined ones can wait. Tap a topic to open it and mark it done, learning or skipped.

Data AnalystRelated roadmapsSQL roadmapPython roadmapExcel roadmapAI & Data Scientist roadmapPrompt Engineering roadmapAI Agents roadmapProve it as you goReading is half of it. This roadmap isbacked by a free, scored Data Analystcertification, a realistic job simulationand daily challenges - evidence arecruiter can verify, not a claim.Get certified - freeJob simulationDaily challengesKnow the job before the toolsThe analyst job in 2026What an analyst actuallydoesAnalyst vs BI vs data scientistvs analytics engineerHow AI changed the jobWhat hiring managers testnowFresher path vs switcherpathMistakes that keep learnersstuckThink like an analystBusiness questions & metricsTurning a vague ask into aquestionThe four kinds of questionsMetrics, KPIs anddefinitionsData grain and uniquenessThe data lifecycleStakeholders, requirementsand pushbackSpreadsheets, still the fastesttoolExcel & Google SheetsCore functions & referencesSUMIFS /COUNTIFSIF / IFS / AND/ ORDates & DATEDIFTRIM & textfunctionsLookups: XLOOKUP &INDEX-MATCHPivot tablesCharts & conditionalformattingPower Query: clean & combineAI inside spreadsheets:Copilot & GeminiSQL is non-negotiableSQL fundamentalsSELECT, WHERE, ORDER BY,LIMITJoins and keysGROUP BY, HAVING andaggregatesNULL handlingDates, strings and CASESQL for analysisSubqueries & CTEsWindow functionsValidating a query: counts &reconciliationCohorts, funnels andretentionEXPLAIN and queryperformanceWriting SQL with AI, safelySQL interview practice setsPython for analysisPython & pandasPython basics for datapandas: load, filter, group,mergeCleaning with pandasNotebooks & reproducibilityPulling data: files, APIs,databasesPlotting: matplotlib,seaborn, plotlyAutomating reports withscriptsAI-assisted coding:generate, review, testData quality: the real 70%Data quality & cleaningProfiling a new datasetMissing data strategiesDuplicates & outliersTypes, formats andstandardisationReconciling against thesource of truthDocumentation & datadictionariesPrivacy: PII, masking andthe DPDP ActStatistics you can defendStatisticsDescriptive statisticsMean, median,modeRange,variance, SDDistributionshapePercentiles &outliersSampling & biasUncertainty & confidenceintervalsHypothesis tests & p-valuesCorrelation vs causationA/B tests & experimentdesignRegression basicsForecasting & time seriesbasicsVisualise it, then tell thestoryDashboards & storytellingChoosing the right chartOne BI tool, properlyBI data modelling: starschema & relationshipsMeasures & DAX basicsKPI design & definitionsData storytelling &executive summariesPresenting and defendingyour numbersAI in BI: Copilot &natural-language queriesThe AI-first analystAI-assisted analysisPrompting for analysis:context, schema, constraintsThe 2026 toolkit: ChatGPT,Claude, Copilot, GeminiVerify everything: thereview loopAI for cleaning &transformation codeDrafting insights andsummaries with AIAnalysing text with LLMs:classify, extract, tagData privacy with AI toolsSmall automations & agentswith LLM APIsLimits: hallucination,leakage, reproducibilityWhere the data livesModern data stack awarenessWarehouses: BigQuery,Snowflake, RedshiftELT and dbt basicsGit & version control foranalystsCloud basics & costawarenessWorking with data engineersMachine learning awarenessSupervised, unsupervised,reinforcementCommon algorithms atintuition levelEvaluating models honestlyWhen not to use MLGet hiredPortfolio & job searchPortfolio projects that getreadCase studies & take-hometestsSQL & analytics interviewprepResume & LinkedIn foranalyst rolesBehavioural & stakeholderstoriesCertifications that actuallymatterProve it: certification &job simulationCommunity, competitions &applying smartlyPortfolio projects that get readRetention teardown: who comes back, and what they areworthMonthly cohorts from a public transactional dataset; end withthe one change you would test next and the number it wouldmove.The data quality incident reportFind a dataset with real problems, quantify how much they movethe headline metric, fix what can be fixed and state whatcannot.One decision dashboard, defendedOne stakeholder, one decision, four charts at most, each withthe action it drives - and a two-minute recorded walkthrough.Kaggle datasetsData Analyst job simulationWhat most learners get wrongCollecting courses instead of finishing one analysis end toend.Learning tools before learning to ask a precise businessquestion.Stopping SQL at joins; window functions and validation arewhere interviews are decided.A portfolio of clean tutorial data - real hiring testsjudgment under mess.Shipping a number without reconciling it against a source oftruth.Pasting AI output into a deliverable without checking thejoin, the grain or the maths.Chasing machine learning before statistics you can defend.Never practising the explanation: the walkthrough is half ofevery interview.Certifications that carry weightOneRoadmap Data Analyst certification - free30 minutes, scored deterministically, and it feeds your TalentGraph so recruiters can verify it in one click.PL-300: Power BI Data Analyst AssociateNamed most often in Indian and Gulf BI job descriptions. Takeit once you can model data - it tests DAX and modelling, notchart-making.Google Data Analytics Professional CertificateThe strongest single credential for a career switcher with noanalytics job history; it forces a portfolio out of you.Get certified - freePL-300 detailsKeep goingGo deeper on the tools this roadmap leans on, or step acrossto the AI and Data Scientist path once the analyst foundationsare in place.SQL roadmapPython roadmapAI & Data Scientist roadmapDaily challenges

Everything on this roadmap

Every stage and topic on the map, in reading order. Click any topic to open it.

The Data Analyst roadmap, topic by topic

Every topic on the map in reading order, with what to learn and the resources we recommend. Main topics are open; open any subtopic to read it. Select a topic name to open it in the topic panel, where you can track your progress.

  1. 1. Know the job before the tools

    The analyst job in 2026

    A data analyst turns business data into decisions people act on: they find the right data, clean it, question it, and explain what it means to someone who has to choose. In 2026 the mechanics are cheaper than ever because AI drafts queries, code and charts, so the value has moved to judgment - knowing which question matters, whether a number can be trusted, and what to do next.

    Fresher: this stage is your map of the territory. Read all of it before touching a tool.
    Experienced: skim for what changed - the hiring signals and the AI shift are new since most people learned the job.

    What an analyst actually does

    Less maths than most expect, more communication than anyone admits. A typical week: clarifying a vague request, pulling data with SQL, cleaning it, checking it against a known total, building one chart that answers the question, and writing three sentences a manager can act on. Dashboards are a by-product; decisions are the product.

    Fresher: notice how little of the week is "learning a tool" - aim your practice at whole cycles, not features.

    Analyst vs BI vs data scientist vs analytics engineer

    Overlapping titles with different centres of gravity. Data analysts answer business questions with SQL, spreadsheets and BI. BI analysts specialise in reporting and dashboards. Analytics engineers own the transformation layer (dbt, warehouses) that analysts query. Data scientists build models and run experiments. Knowing the map aims your learning and your job search - and stops you applying to roles that want a different person.

    How AI changed the job

    AI

    Since 2023, assistants write first drafts of SQL, pandas code, DAX and even the insight paragraph. That did not remove analysts; it removed the value of typing. What employers now pay for is the person who can specify the question precisely, check the draft against the schema and the totals, notice a wrong join or a silent NULL, and take responsibility for the number. This roadmap treats AI as a tool you must use well, not as a shortcut past understanding.

    Both tracks: every AI-tagged node on this map pairs a capability with the check that makes it safe to use.

    What hiring managers test now

    Hiring

    Screening rounds in 2026 look alike across companies: a timed SQL test (joins, aggregation, window functions), a take-home case study on a messy dataset, a walkthrough of one portfolio project, and behavioural questions about a stakeholder who disagreed with you. Tools are rarely the filter; reasoning under mess and clear explanation are. Build your preparation around those four moments.

    Fresher: your portfolio replaces experience - make one project deep rather than five shallow.
    Experienced: expect the case study to probe how you defined the metric and validated the result, not just the chart.

    Fresher path vs switcher path

    Optional on the Experienced track

    A fresher builds credibility from scratch: certification, one deep portfolio project, SQL fluency, and the ability to explain both. A career switcher already has a domain - finance, operations, marketing, healthcare - and should lead with it: analyse the data of the field you know, because domain sense is the part of analytics that cannot be taught in a week. Both paths converge on the same interviews.

    Mistakes that keep learners stuck

    Read first

    The pattern repeats in every cohort: collecting courses without finishing one end-to-end analysis; learning tools before learning to ask a precise question; stopping SQL at joins; a portfolio of clean tutorial datasets; shipping a number without reconciling it; pasting AI output into a deliverable unchecked; chasing machine learning before statistics; never rehearsing the explanation. Every one of these has a node on this map marked "Missed". Do those first.

  2. 2. Think like an analyst

    Business questions & metrics

    Every analysis starts with a question a business actually has, translated into a metric with a precise definition and a known grain. Skipping this step is why so many dashboards exist and so few decisions change. This topic is tool-free on purpose: it is the part of the job that AI cannot do for you.

    Turning a vague ask into a question

    "Can you look at churn?" is not a question. "What share of customers who joined in Q1 made no purchase in the following 90 days, and how does that compare with Q4?" is. Practise rewriting requests until they name the population, the metric, the time window and the comparison. Then confirm the rewrite with the requester before you open a query editor - it saves days.

    Experienced: this is also how you push back on a bad request without saying no.

    The four kinds of questions

    Every request is one of four: what happened (descriptive), why did it happen (diagnostic), what will happen (predictive), or what should we do (prescriptive). Identifying which one you are being asked decides your entire approach - a descriptive question needs a clean aggregate, a diagnostic one needs segmentation and comparison, a predictive one needs a model and honesty about uncertainty. Most analyst work is the first two, done well.

    Metrics, KPIs and definitions

    A metric is a definition, not a number: numerator, denominator, time window, filters, and the exact table it comes from. "Active users" means nothing until you say active how, over what period, excluding whom. Write definitions down where the whole team can see them, and never let two dashboards report two different values for the same name.

    Data grain and uniqueness

    Missed

    Grain is what one row represents: one order, one order line, one customer per month. Almost every wrong total comes from joining or aggregating at the wrong grain - a join that fans out one order into three lines silently triples revenue. Before any analysis, state the grain of every table in one sentence and check it with a uniqueness query. This single habit prevents the most common mistake in interviews and on the job.

    The data lifecycle

    Collection, cleaning, exploration, analysis, visualisation, decision - and then measurement of whether the decision worked. Knowing where you are in the cycle keeps a project from becoming aimless chart-making, and tells you what "done" means at each step. Most learner projects skip the last two steps; most employers only care about them.

    Stakeholders, requirements and pushback

    Analysts work for people who are busy, sure of their hypothesis, and often wrong about what data exists. Learn to run a fifteen-minute requirements conversation, to write back what you heard, to say "the data cannot answer that, but it can answer this", and to deliver an inconvenient result without a fight. Interviews probe this directly with behavioural questions.

    Fresher: the job simulation on this roadmap puts you in this conversation with a realistic manager.

  3. 3. Spreadsheets, still the fastest tool

    Excel & Google Sheets

    Spreadsheets are still the fastest way from a file a stakeholder emailed you to a first answer, and the tool every business person already trusts. Master the analytical core - lookups, pivots, Power Query - rather than every function. Then let AI write the awkward formulas while you check them.

    Fresher: do the functions box; interviews for entry roles still ask them.
    Experienced: skip straight to Power Query and the AI features.

    Core functions & references

    Optional on the Experienced track

    Absolute versus relative references, named ranges, and the everyday aggregates - SUM, AVERAGE, COUNT, MIN and MAX - that everything else builds on. Get comfortable reading a formula someone else wrote; you will inherit far more workbooks than you create.

    SUMIFS / COUNTIFS

    Optional on the Experienced track

    Conditional aggregation in a sheet: total sales for one region and one month, count of orders above a threshold. SUMIFS and COUNTIFS are the spreadsheet twins of SQL's GROUP BY with a WHERE, and the fastest way to answer a business question without a pivot.

    IF / IFS / AND / OR

    Optional on the Experienced track

    Business rules inside a sheet: flag late deliveries, bucket customers into tiers, mark exceptions. Nest sparingly - an IFS or a lookup table is usually clearer than five nested IFs, and clearer means checkable.

    Dates & DATEDIF

    Optional on the Experienced track

    Dates are numbers in disguise, which is why they break. Learn DATE, EOMONTH, DATEDIF and how to spot a date stored as text. Days between order and delivery, months since signup and fiscal periods are everyday analyst arithmetic.

    TRIM & text functions

    Optional on the Experienced track

    TRIM, SUBSTITUTE, UPPER/LOWER/PROPER, LEFT/RIGHT/MID and TEXTSPLIT standardise human-typed chaos into analysable columns. Half of "the numbers don't match" in a spreadsheet is a trailing space or a mixed-case category.

    Lookups: XLOOKUP & INDEX-MATCH

    XLOOKUP (and the legacy VLOOKUP you will meet in old workbooks) plus INDEX-MATCH are how spreadsheets join data across sheets. Learn exact match, what happens on no match, and why a lookup that returns the first duplicate silently corrupts a report. A guaranteed interview topic for entry-level roles.

    Pivot tables

    Group, aggregate and slice thousands of rows in seconds - the fastest route from raw data to a first answer. Learn rows versus columns versus values, how to add a calculated field, and how to spot when the source range silently excluded new rows.

    Charts & conditional formatting

    Optional on the Experienced track

    Quick trend, comparison and share charts directly from ranges and pivots - analysis at the speed of conversation - plus conditional formatting to make exceptions jump out of a table. Good enough for a meeting; not a substitute for a BI tool.

    Power Query: clean & combine

    Power Query (Excel and Power BI) records cleaning steps - remove blanks, split columns, unpivot, merge files - as a repeatable script, so next month's file cleans itself. It is the bridge between spreadsheet habits and proper data pipelines, and the most under-learned Excel feature in analyst interviews.

    AI inside spreadsheets: Copilot & Gemini

    AI

    Copilot in Excel and Gemini in Sheets write formulas, suggest pivots and explain existing sheets. Use them to draft, then verify: check a formula on three rows by hand, confirm a pivot's total against the raw column, and never accept a generated lookup without testing the no-match case. The productivity is real; so is the confident wrong answer.

  4. 4. SQL is non-negotiable

    SQL fundamentals

    SQL is the one skill every analyst interview tests and every analyst job uses daily. Pull your own data instead of waiting on someone else, and learn to reason about row counts before you run a query. This topic is the syntax; the next one is analysis.

    Experienced: the basics node is outlined for you - jump to NULLs and dates, which still trip people with years of experience.

    SELECT, WHERE, ORDER BY, LIMIT

    Optional on the Experienced track

    Reading exactly the rows you need: filtering with WHERE, sorting, limiting, expressions and aliases. Learn operator precedence with AND/OR and always parenthesise mixed conditions - a missing bracket is a classic silent bug.

    Joins and keys

    INNER, LEFT and anti-joins, and reasoning about row counts before you run the query. Identify the join key, check it is unique on at least one side, and predict whether rows will multiply or disappear. Most interview SQL is join logic in disguise, and most wrong revenue numbers are a fan-out.

    GROUP BY, HAVING and aggregates

    Turning transaction tables into the metrics a business runs on: COUNT, SUM, AVG, MIN, MAX, COUNT(DISTINCT), grouping by one or several columns, filtering groups with HAVING, and conditional aggregation with CASE inside SUM. Know why HAVING is not WHERE.

    NULL handling

    NULL is not zero, not empty and not equal to anything, including itself. It changes the result of comparisons, drops out of COUNT(column) but not COUNT(*), disappears from averages, and turns arithmetic into NULL. Learn IS NULL, COALESCE and NULLIF, and check every LEFT JOIN for the NULLs it creates. Still the most common source of a wrong answer.

    Dates, strings and CASE

    Truncating dates to month, extracting the year, computing intervals, casting text to dates, cleaning strings with TRIM and LOWER, and bucketing values with CASE. Dialects differ here more than anywhere - know your database's date functions rather than memorising one syntax.

    SQL for analysis

    Beyond syntax: the SQL that answers real questions and survives review. CTEs make multi-step logic readable, window functions rank and run totals, cohort queries measure retention, and validation queries prove the number before anyone sees it. This is the layer that separates juniors from mid-levels in screening rounds.

    Subqueries & CTEs

    WITH clauses that make multi-step logic readable, correlated subqueries, and knowing when each beats a join. Name every CTE for what its rows are (one row per customer per month), and you will catch grain mistakes as you write.

    Window functions

    ROW_NUMBER, RANK, LAG/LEAD, running totals and moving averages over PARTITION BY ... ORDER BY. They answer "first order per customer", "change from last month" and "top three per region" without self-joins - and they appear in almost every analyst SQL test.

    Validating a query: counts & reconciliation

    Missed

    Before a number leaves your hands: does the row count match the grain you expected, does the total reconcile with a known figure (finance's revenue, last month's report), does one spot-checked entity look right end to end, and did a filter or a NULL drop rows you meant to keep? Write these checks as queries and keep them next to the analysis. Most learners never do this; every good analyst does it every time.

    Cohorts, funnels and retention

    The three analyses product and growth teams ask for most. A cohort groups users by when they started; retention measures how many are still active n periods later; a funnel counts how many reach each step in order. All three are grain exercises - one row per user per period - built with date truncation, joins and window functions.

    EXPLAIN and query performance

    Optional on the Fresher track

    When a query takes minutes, read its plan. Learn to spot a full scan on a large table, a join without a usable key, and a filter applied after an expensive step. On warehouses, performance is also cost - know what partition pruning and column selection save.

    Experienced: expected of you; Fresher: awareness is enough for now.

    Writing SQL with AI, safely

    AI

    Give the assistant the schema (tables, keys, grain), the metric definition and the dialect, and ask for the query plus the assumptions it made. Then check the join keys, the grain of the result, the NULL handling and the date boundaries, and run a validation query against a known total. Treat generated SQL as a junior's draft: fast, useful, and yours to review.

    SQL interview practice sets

    Hiring

    Timed practice beats reading. Work the LeetCode SQL 50, then StrataScratch or DataLemur questions from real companies, and rehearse explaining your query out loud. Aim to reason about row counts before running - that is what interviewers listen for.

  5. 5. Python for analysis

    Python & pandas

    Python is how an analyst automates what spreadsheets cannot and unlocks the PyData stack: pandas for data, matplotlib and plotly for charts, notebooks for reproducible work, and small scripts that replace a monthly copy-paste ritual. In 2026 you will write much of it with an AI assistant - which makes reading and testing code the skill that matters.

    Fresher: learn just enough Python to be dangerous in a notebook, then live in pandas.
    Experienced: skip the basics; focus on cleaning, automation and AI-assisted coding.

    Python basics for data

    Optional on the Experienced track

    Variables, lists, dicts, loops, functions, files and the standard library - enough of the language to read and fix code, not to build software. R is the main alternative for statistics-heavy roles; pick one and learn it well.

    pandas: load, filter, group, merge

    DataFrames are spreadsheets with superpowers: read a CSV or a database table, filter rows, create columns, group and aggregate, merge tables and reshape wide to long. Learn the mental mapping from SQL - merge is join, groupby is GROUP BY - and you can move between the two without friction.

    Cleaning with pandas

    Detect and handle missing values, drop or flag duplicates, fix types (dates stored as text, numbers with commas), standardise categories and reshape. Write the cleaning as a function you can re-run, and print row counts before and after each step so nothing disappears silently.

    Notebooks & reproducibility

    Jupyter is the analyst's workbench: code, output and narrative in one document. Reproducibility means someone else can run it top to bottom and get the same result - fixed random seeds, pinned package versions, data loaded from a stated source, no hidden state from cells run out of order. Restart and run all before you share.

    Pulling data: files, APIs, databases

    Where analyst data actually comes from: CSV and Excel exports, REST APIs with pagination and rate limits, SQL databases through a connector, and - carefully and legally - web pages. Learn one clean pattern for each and the trade-offs: files are simple but stale, APIs are live but fragile, databases are authoritative but need access.

    Plotting: matplotlib, seaborn, plotly

    matplotlib for control, seaborn for statistical plots in one line, plotly for interactive charts you can hand to a stakeholder. Exploration plots exist to expose bad data before it fools you - a histogram and a scatter per key column is the cheapest data-quality check you will ever run.

    Automating reports with scripts

    Optional on the Fresher track

    Anything you do monthly by hand should become a script: pull, clean, compute, chart, export to Excel or a slide, send. Learn to parameterise by date, log what ran, and fail loudly when a source changes shape. Scheduling (cron, GitHub Actions, or a warehouse scheduler) turns a script into a service.

    Experienced: this is where you get your Fridays back.

    AI-assisted coding: generate, review, test

    AI

    Copilot, ChatGPT and Claude write pandas faster than you can type. Use them for the first draft and for explaining unfamiliar code; then read every line, run it on a small sample with a known answer, and add an assertion for the row count and the total. The skill employers now pay for is reviewing generated code, not producing it.

  6. 6. Data quality: the real 70%

    Data quality & cleaning

    The unglamorous 70% of the job. Real data arrives with duplicates, impossible dates, silent NULLs, three spellings of the same city and a total that does not match finance. Cleaning is not a chore before analysis - it is analysis, and it is where judgment shows. Document every choice, because it changes your conclusions.

    Profiling a new dataset

    Before analysing, describe: row count, grain, column types, missing share per column, distinct values of every category, min and max of every number and date, and a few example rows. Ten minutes of profiling catches most surprises and gives you the questions to ask the data owner.

    Missing data strategies

    Detect NaNs and their disguises (empty strings, zeros, 9999, "N/A"), then decide: drop, fill, flag or model - and say which in the write-up. Missingness is often informative: customers without a phone number behave differently. Never fill silently with a mean and move on.

    Duplicates & outliers

    Deduplicate on the key that defines the grain, not on every column; then check what the duplicates were telling you. For outliers, use IQR or z-scores to find them and judgment to decide: is this point an error, or the most interesting row in the file? Removing outliers changes averages - state that you did.

    Types, formats and standardisation

    Cast text to dates and numbers, unify units and currencies, normalise categories (Delhi, delhi, New Delhi), trim whitespace, and reshape wide to long. A lookup table of canonical values that the whole team shares beats a hundred one-off replacements.

    Reconciling against the source of truth

    Missed

    Every headline number needs an anchor: revenue against finance, orders against the operational system, users against the product team's dashboard. Reconcile totals, explain differences to the rupee or the row, and write the reconciliation down. It is the difference between "my query says" and "the number is". Most learners never learn this because tutorials have no second source.

    Documentation & data dictionaries

    A data dictionary names each table's grain and each column's meaning, type, allowed values and source. A short README states where data came from, what was cleaned and what is still wrong. Interviewers read these to judge how you think; teammates read them to avoid repeating your work.

    Privacy: PII, masking and the DPDP Act

    2026

    Names, phones, emails, Aadhaar numbers and locations are personal data. Know what you may collect and keep, mask or hash it before analysis, aggregate before sharing, and never paste raw customer rows into an AI tool. India's Digital Personal Data Protection Act (2023) and the GDPR make this a legal duty, and hiring managers increasingly ask about it.

  7. 7. Statistics you can defend

    Statistics

    Say "this difference is real" and defend it. Analysts need working statistics, not proofs: describe a distribution honestly, know how sampling misleads, put uncertainty on an estimate, read a test without being fooled by noise, and design a fair experiment. Skip this stage and every dashboard you build will eventually lie to someone.

    Descriptive statistics

    The numbers that summarise a column: centre, spread, shape and position. Always look at the distribution before quoting the average - a mean with a few whales in it describes nobody.

    Mean, median, mode

    Mean for symmetric data, median when a few large values distort the picture (income, order value, response time), mode for categories. Report the one that describes a typical case, and say why.

    Range, variance, SD

    A metric's average means little until you know its spread. Range, interquartile range, variance and standard deviation tell you whether the average is representative and whether a change is large relative to normal noise.

    Distribution shape

    Skewed, bimodal, heavy-tailed or roughly normal - the shape decides which summary and which test are honest. Plot a histogram before you compute anything; normality is the silent assumption behind many tests.

    Percentiles & outliers

    P50, P90 and P99 describe experience better than averages for anything with a long tail - delivery times, page loads, ticket resolution. Percentiles also define outliers (the IQR rule) and service targets.

    Sampling & bias

    A survey of customers who answered is not a sample of customers. Learn selection bias, survivorship bias and non-response, why sample size shrinks noise but not bias, and how to state who your data actually represents. Most wrong business conclusions are sampling problems, not maths problems.

    Uncertainty & confidence intervals

    An estimate from a sample is a range, not a point. Confidence intervals put honest bounds on a conversion rate or an average, and stop a 2% "improvement" measured on 50 users from becoming a headline. Report intervals whenever the sample is small or the decision is expensive.

    Hypothesis tests & p-values

    Null hypotheses, p-values and the difference between statistically significant and practically important. Know what a p-value is not (the probability the result is due to chance), why peeking at results inflates false positives, and how to explain a test to a manager in one sentence.

    Correlation vs causation

    Measuring relationships, spotting confounders, and the honesty to say "correlated" when you cannot say "caused". Ice cream and drownings both rise in summer. Before claiming a driver, ask what else changed and whether an experiment could settle it.

    A/B tests & experiment design

    Randomisation, a pre-registered metric, a sample size you computed before starting, a fixed duration, and one comparison - that is a fair test. Learn the common failures (peeking, multiple metrics, novelty effects, segment fishing) because you will be asked to read a test someone else ran badly.

    Regression basics

    Linear regression as the workhorse of "what drives this metric?" - coefficients as effect sizes, fit as how much is explained, residuals as a sanity check. Logistic regression for yes/no outcomes. Enough to read a data scientist's output and challenge it.

    Forecasting & time series basics

    Optional on the Fresher track

    Trend, seasonality and noise; moving averages and exponential smoothing; why a forecast needs an interval and a horizon; and how to judge one against a naive baseline. Demand, revenue and capacity questions all arrive as forecasts.

    Experienced: expected in most senior analyst roles. Fresher: know the vocabulary.

  8. 8. Visualise it, then tell the story

    Dashboards & storytelling

    Ship a dashboard a manager checks every Monday, and write the paragraph that tells them what to do about it. Visualisation is a decision-support tool, not decoration: the right chart for the question, a data model that does not double count, KPIs defined once, and a story that leads with the answer.

    Choosing the right chart

    Bar for comparison, line for trend, scatter for relationship, histogram for distribution, heatmap for two-way patterns, funnel for stages, and (rarely) pie for a share of a whole - matching the chart to the question, and knowing when a table beats them all. Start every chart with a title that states the finding.

    One BI tool, properly

    Pick Power BI, Tableau or Looker Studio - whichever the jobs you want name - and go deep: connecting data, a proper model, calculated fields, filters and slicers, drill-through, scheduled refresh and sharing. One tool learned properly transfers; three learned superficially do not.

    Fresher in India or the Gulf: Power BI is named most often in job descriptions.

    BI data modelling: star schema & relationships

    Optional on the Fresher track

    Facts (orders, events) in the middle, dimensions (customers, products, dates) around them, one-to-many relationships flowing the right way, and a date table. Get this wrong and every total in the report double counts or filters incorrectly. PL-300 tests this more than charts.

    Experienced: core - most dashboard bugs are model bugs.

    Measures & DAX basics

    Optional on the Fresher track

    Measures compute at query time in the report's filter context: total sales, sales last year, share of total, running totals. Learn CALCULATE, filter context, and time intelligence - and why a calculated column is usually the wrong answer. Tableau's LOD expressions and Looker's measures cover the same ground.

    KPI design & definitions

    Define metrics precisely - numerator, denominator, time window, filters - so two people never report two numbers for the same KPI. Put the definition on the dashboard, version it, and pair every KPI with the decision it informs. A KPI nobody acts on is a chart.

    Data storytelling & executive summaries

    Overview first, drill-down second; every view answers a question someone actually asks. In writing: lead with the answer, then the evidence, then the caveat, then the recommendation - in that order, in under 150 words. Executives read the first sentence; make it the finding.

    Presenting and defending your numbers

    Hiring

    Every interview and every review has the moment when someone asks "are you sure?". Be able to say where the data came from, what the grain is, what you excluded and why, what the number was last period, and what would change your conclusion. Rehearse a two-minute walkthrough of your best project until it is boring.

    AI in BI: Copilot & natural-language queries

    AI

    Copilot in Power BI, Tableau Pulse and similar features let users ask questions in plain language and get charts, DAX and summaries. They are only as good as the model and the metric definitions underneath - which makes your modelling work more important, not less. Learn what they do well (drafting, explaining) and where they mislead (ambiguous metrics, wrong grain).

  9. 9. The AI-first analyst

    AI-assisted analysis

    AI

    The 2026 analyst works with an assistant open: it drafts SQL and pandas, explains unfamiliar code, summarises a table, classifies free text and writes the first version of the insight. The analyst's job is to give it the right context, verify everything it returns, and take responsibility for the result. This stage is the difference between an analyst AI replaces and one AI multiplies.

    Prompting for analysis: context, schema, constraints

    A good analysis prompt states the goal, the tables with their keys and grain, the metric definition, the dialect or library, the output format, and what to do when unsure. Ask for assumptions to be listed. Iterate: paste the error or the wrong total back and ask for the fix. Vague prompts get confident, wrong answers.

    The 2026 toolkit: ChatGPT, Claude, Copilot, Gemini

    ChatGPT's data analysis mode runs Python on an uploaded file; Claude reasons well over long schemas and documents; GitHub Copilot lives in your editor and notebook; Copilot in Excel and Power BI and Gemini in Sheets sit inside the tools you already use. Learn one general assistant well and the in-tool features for your stack, and know what each is allowed to see.

    Verify everything: the review loop

    Missed

    Generated SQL, code and text are drafts. For every AI output: check the join keys and grain, run it on a sample with a known answer, reconcile the total, read the maths in any summary, and confirm every quoted number appears in the data. Keep the checks as code next to the analysis. Most people skip this; it is the single habit that lets you use AI aggressively without getting burned.

    AI for cleaning & transformation code

    Describe the mess - three date formats, categories with typos, a wide table that should be long - and let the assistant write the pandas, Power Query or SQL. Then test on the ugliest rows you can find, print row counts before and after, and read the regular expressions it produced. Cleaning code is where generated code most often looks right and is wrong.

    Drafting insights and summaries with AI

    Give the model the final table and the audience, ask for a three-sentence summary that leads with the finding, and then edit: remove any claim the data does not support, add the caveat, name the action. AI is good at fluent prose and bad at knowing what matters to this manager - that part stays yours.

    Analysing text with LLMs: classify, extract, tag

    Support tickets, reviews, survey comments and call notes were unanalysable at scale until recently. Now an LLM can classify sentiment, extract the product mentioned, tag the reason for churn, and return structured rows you can count. Design the label set first, sample and hand-check a hundred results for accuracy, and report that accuracy with the numbers.

    Data privacy with AI tools

    Never paste raw customer rows, credentials or unreleased financials into a consumer AI tool. Know your company's approved tools and their data-retention terms, aggregate or mask before you share, and prefer enterprise deployments that do not train on your data. A privacy incident ends careers faster than a wrong number.

    Small automations & agents with LLM APIs

    Optional on the Fresher track

    With a few dozen lines of Python and an LLM API you can classify incoming tickets nightly, draft a weekly metrics note, or answer common data questions from a schema. Start with one narrow task, log every input and output, keep a human in the loop for anything that leaves the team, and measure accuracy before you trust it.

    Experienced: this is fast becoming the analyst's edge. Fresher: awareness now, build later.

    Limits: hallucination, leakage, reproducibility

    Models invent plausible numbers, functions and citations; they may have seen the answer to a public dataset; and the same prompt can return different code tomorrow. Treat outputs as untrusted, pin versions and seeds where you can, save the exact prompt and result with the analysis, and never report a figure you cannot reproduce from the data yourself.

  10. 10. Where the data lives

    Modern data stack awareness

    Optional on the Fresher track

    Analysts increasingly query a cloud warehouse fed by ELT tools and modelled with dbt, with everything in Git. You do not need to build it, but you need to speak it: where a table comes from, why it refreshed late, what a query costs, and how to ask an engineer for a change.

    Experienced: core - most mid-level roles expect it. Fresher: know the words; Git is the one to actually learn.

    Warehouses: BigQuery, Snowflake, Redshift

    Optional on the Fresher track

    Columnar, cloud-scale databases built for analytics. Learn what makes them different from a transactional database (columns, partitions, separation of storage and compute), why SELECT * costs money, and how to use partition filters. Most analyst SQL in 2026 runs on one of these.

    ELT and dbt basics

    Optional on the Fresher track

    Extract-load-transform: raw data lands in the warehouse first and is transformed there with SQL models. dbt organises those models, tests them and documents them. Understanding it tells you which table to trust, and writing a small dbt model is the natural next step for an analyst who owns metric definitions.

    Git & version control for analysts

    Commit your SQL, notebooks and dbt models; branch for a change; open a pull request so someone reviews the metric definition before it ships. Git is how analysis stops being a file called final_v3_REAL.xlsx, and it is expected in most teams now.

    Cloud basics & cost awareness

    Optional on the Fresher track

    Storage buckets, IAM permissions, service accounts and the bill. Know how to request access properly, why a scheduled query that scans a terabyte a day matters, and how to read a cost dashboard. Analysts who understand cost get more freedom.

    Working with data engineers

    Optional on the Fresher track

    Write a good ticket: the table, the column, the grain, the example rows that look wrong, the business impact and a deadline. Learn the pipeline's refresh schedule and failure alerts. Engineers help most the people who bring specifics.

    Machine learning awareness

    Optional on the Fresher track

    Speak the language of the team next door. An analyst should recognise which business problems map to which model types, read an evaluation honestly, and know when a simpler method beats a model. You are not training models here - you are learning enough to ask the embarrassing question.

    Fresher: outlined on purpose - finish statistics first.

    Supervised, unsupervised, reinforcement

    Optional on the Fresher track

    Supervised learning predicts a label from examples (churn, fraud, price); unsupervised finds structure without labels (segments, anomalies); reinforcement learns from rewards (rare in analytics). Which business problems map to which, and where the analyst hands off to a data scientist.

    Common algorithms at intuition level

    Optional on the Fresher track

    Linear and logistic regression, decision trees and random forests, k-nearest neighbours, naive Bayes and k-means - the recognisable cast of everyday ML. Know what each is for and what it assumes; leave the maths to the people who tune them.

    Evaluating models honestly

    Optional on the Fresher track

    Train/test splits, why accuracy lies on imbalanced data, precision versus recall, and leakage - when a feature secretly contains the answer. Enough to ask "how was this validated, and on what period?" before a model's output reaches a dashboard.

    When not to use ML

    Missed

    Most business questions are answered by a good GROUP BY, a chart and a conversation. A model costs time, needs maintenance, and hides its reasoning. Reach for it when the pattern is genuinely complex, the data is plentiful, and the decision repeats. Saying "we do not need a model for this" is a senior skill.

    Big data & Spark, concept level

    Not on the Fresher track · Optional on the Experienced track

    When data outgrows one machine: distributed storage and processing with Spark, and the Hadoop lineage it replaced. Warehouses hide most of this from analysts now; know the concepts, not the cluster tuning.

    Deep learning (optional)

    Not on the Fresher track · Optional on the Experienced track

    Neural networks, CNNs and transformers with PyTorch or TensorFlow - optional territory for analysts, useful vocabulary in AI-heavy teams. The LLMs on this roadmap are deep learning; you use them without training them.

  11. 11. Get hired

    Portfolio & job search

    Hiring

    Turn analyses into interviews. Two or three end-to-end projects on messy real data beat ten tutorial notebooks; a rehearsed walkthrough beats a long resume; verified proof beats a claim. This stage is about presenting what you can do to the people who decide.

    Fresher: your portfolio is your experience. Experienced: lead with impact and domain, and refresh the SQL test muscle.

    Portfolio projects that get read

    Pick messy, public, real data. Ask one business question, clean and validate, answer it, and end with a recommendation and the number it would move. Publish the notebook or dashboard, a README with the data dictionary and the cleaning decisions, and a two-minute recorded walkthrough. The three projects in the box at the bottom of this map are designed for exactly this.

    Case studies & take-home tests

    You get a dataset and a vague question and 48 hours. Winners restate the question precisely, profile and validate the data first, answer with two or three charts and a recommendation, list what they would do with more time, and keep the whole thing short. Losers build twelve charts and never say what to do.

    SQL & analytics interview prep

    Expect a live SQL round (joins, aggregation, window functions, a NULL trap), a metrics question ("how would you measure engagement?"), a probability or statistics question, and "tell me about a time your analysis was wrong". Practise all four out loud; the reasoning you narrate is what is scored.

    Resume & LinkedIn for analyst roles

    Quantified bullets (what you analysed, what changed, by how much), tools named where they were used, a link to the portfolio, and nothing that a screening question would expose as inflated. Match the language of the job description without keyword stuffing; applicant tracking systems rank on it. One page for freshers.

    Behavioural & stakeholder stories

    Prepare five short stories in situation-action-result form: a stakeholder who disagreed, a number you got wrong and fixed, a vague request you sharpened, a deadline you protected, a finding nobody wanted to hear. Interviewers are checking judgment and communication - the parts of the job tools cannot do.

    Certifications that actually matter

    Certifications open doors when they are recognised and cheap relative to the signal. The OneRoadmap Data Analyst certification is free, scored and verifiable. PL-300 is named most often in Indian and Gulf BI roles; take it once you can model data. The Google Data Analytics certificate suits switchers with no analytics history. Tableau Desktop Specialist only if the roles you want name Tableau.

    Prove it: certification & job simulation

    Take the free OneRoadmap Data Analyst certification and the Data Analyst job simulation. Both are scored, both produce a verification link a recruiter can check in one click, and both feed the Talent Graph that ranks you for employers on this platform. Claims are cheap; this is not.

    Switching in: leverage your domain

    Not on the Fresher track

    If you come from finance, operations, marketing, sales or healthcare, you already have what freshers lack: you know which questions matter and what the numbers should look like. Build your portfolio on your own field's data, apply to analyst roles in that field first, and frame your experience as analysis you were already doing without the title.

    Community, competitions & applying smartly

    Optional on the Experienced track

    Learning in public compounds: share a project write-up, enter a Kaggle competition for the feedback, and join analytics communities where hiring managers actually post. Apply in focused batches to roles that match your portfolio, and follow up with a specific comment about the company's data - it works far more often than volume.

Prove it as you go

Reading is half of it. This roadmap is backed by a free, scored Data Analyst certification, a realistic job simulation and daily challenges - evidence a recruiter can verify, not a claim.

Portfolio projects that get read

Retention teardown: who comes back, and what they are worth

Monthly cohorts from a public transactional dataset; end with the one change you would test next and the number it would move.

The data quality incident report

Find a dataset with real problems, quantify how much they move the headline metric, fix what can be fixed and state what cannot.

One decision dashboard, defended

One stakeholder, one decision, four charts at most, each with the action it drives - and a two-minute recorded walkthrough.

What most learners get wrong

Collecting courses instead of finishing one analysis end to end.

Learning tools before learning to ask a precise business question.

Stopping SQL at joins; window functions and validation are where interviews are decided.

A portfolio of clean tutorial data - real hiring tests judgment under mess.

Shipping a number without reconciling it against a source of truth.

Pasting AI output into a deliverable without checking the join, the grain or the maths.

Chasing machine learning before statistics you can defend.

Never practising the explanation: the walkthrough is half of every interview.

Certifications that carry weight

OneRoadmap Data Analyst certification - free

30 minutes, scored deterministically, and it feeds your Talent Graph so recruiters can verify it in one click.

PL-300: Power BI Data Analyst Associate

Named most often in Indian and Gulf BI job descriptions. Take it once you can model data - it tests DAX and modelling, not chart-making.

Google Data Analytics Professional Certificate

The strongest single credential for a career switcher with no analytics job history; it forces a portfolio out of you.

Keep going

Go deeper on the tools this roadmap leans on, or step across to the AI and Data Scientist path once the analyst foundations are in place.

Frequently asked questions

Is data analysis a good career choice in 2026?

Yes. Every company produces data and needs someone to turn it into decisions, so analyst roles exist in every industry - finance, e-commerce, healthcare, SaaS, consulting, government. AI has not removed the job; it has raised the bar. Writing a query is cheap now, so the value has moved to defining the right metric, checking that a number is true and explaining what to do about it. Analysts who work that way are in more demand, not less, and the role leads naturally into analytics engineering, product analytics or data science.

Are data analysts well paid?

Reasonably, and pay rises fast with SQL depth, BI modelling and domain knowledge. In India, freshers at product and consulting firms typically start around ₹4-8 LPA (service companies pay less), analysts with three to five years earn roughly ₹10-20 LPA, and senior or lead analysts more. In the US the common range is about $65k-$110k. Treat these as ranges to verify on current job boards for your city, not promises - the fastest way to move up the range is a portfolio that shows judgment, not a list of tools.

What qualifications do I need to become a data analyst?

No specific degree. Employers hire on demonstrated skills: SQL you can defend in a live test, one BI tool learned properly, working statistics, and a portfolio of two or three projects on messy real data with a written recommendation. Any degree helps with screening; a quantitative or business one helps a little more. Certifications open doors for freshers and switchers - the free OneRoadmap Data Analyst certification, Google's Data Analytics certificate and Microsoft's PL-300 are the ones job descriptions actually name.

Can I teach myself data analysis, and how long does it take?

Yes - most working analysts did. Following this roadmap for about two hours a day, a fresher can be interview-ready in four to six months; a career switcher who already knows a domain is usually faster. What slows people down is collecting courses instead of finishing one analysis end to end. Pick one dataset, ask one business question, clean, validate, answer and present it, then repeat. Three of those beat any number of tutorials.

Do I need to code? Is Python required?

SQL is non-negotiable - every analyst interview tests it and every analyst job uses it. Python is strongly recommended in 2026: most job descriptions list it, it is what you automate with, and AI assistants have made it far cheaper to learn because they write the first draft while you learn to read and check it. Some companies still hire on Excel plus a BI tool alone, but that is a shrinking door.

Is data analysis an IT job?

Not in the traditional sense. Analysts sit between the business and the data: they usually report into an analytics, product, finance or operations team rather than IT, and their output is a decision, not a system. The work is technical - SQL, Python, BI tools, statistics - but it is judged by whether the business acted on the answer, which is why communication is on this roadmap next to the tools.

How is a data analyst different from a data engineer, a business analyst or a data scientist?

A data engineer builds and runs the pipelines and warehouses the analyst queries. A business analyst focuses on requirements and process, often with lighter data skills. A data scientist builds models and runs experiments. A data analyst answers business questions from existing data with SQL, spreadsheets, statistics and BI, and increasingly with AI assistants - and hands off to the others where the problem needs a pipeline or a model.

Will AI replace data analysts?

AI replaces the typing, not the judgment. Assistants already draft SQL, pandas code, DAX and summaries, so the parts of the job that were mostly mechanics are worth less. The parts that remain - specifying the question, knowing the grain of the data, catching a wrong join or a silent NULL, reconciling a number against a source of truth, and defending the result to someone who disagrees - are worth more. This roadmap treats AI as a tool you must use well, with a verification step on every AI-tagged topic.

Which BI tool should I learn: Power BI, Tableau or Looker Studio?

Learn the one the jobs you want name, and learn it properly. In India and the Gulf that is most often Power BI; in the US it is a mix of Power BI and Tableau; Looker Studio is free and good for practice and for Google-heavy companies. The skills that transfer - a proper data model, precise metric definitions, the right chart for the question - matter more than the tool. One tool learned deeply beats three learned superficially.

How do I get hired as a fresher with no experience?

Make the portfolio do the work experience would: two or three projects on messy public data, each ending in a recommendation and a two-minute recorded walkthrough. Get SQL fluent with timed practice sets. Take a free, verifiable certification and a job simulation so recruiters can check your skills in one click. Apply in focused batches to roles that match your projects, and use internships, freelance analysis or your current employer's data to get the first real dataset on your resume.