AI-Augmented Data Analysis
AI-assisted analysis, natural-language BI, agent architectures, context management, accuracy risks, and the changing analyst workflow.
32.6%
Best tweets about Data Science
Browse the best tweets about data science, including analysis, experimentation, statistics, datasets, visualization, careers, and practical workflows.
Useful data science methods, experiments, statistics, analysis, visualization, tooling, career lessons, and real project results.
Original Xholic analysis
The supplied dataset is led by practical learning resources, analysis workflow guidance, and AI-assisted analytics. Learning-oriented posts occupy all five supplied score outliers, while workflow and AI posts emphasize problem framing, context, and scrutiny of AI-generated results. [2014591613582938276, 2070446993105924239, 2040103282106929176, 2054631157191598294]
65.2% of posts
All-time engagement
43.5% of posts
Published in 90 days
Conversation map
AI-assisted analysis, natural-language BI, agent architectures, context management, accuracy risks, and the changing analyst workflow.
32.6%
Free courses, textbooks, roadmaps, coding practice, and project-based paths for learning analytics, statistics, Python, ML, R, and AI.
32.6%
Data cleaning, exploration, metric definition, problem framing, domain context, and turning raw data into trustworthy insight.
28.3%
Machine-learning practice, model development lifecycles, forecasting, evaluation, deployment, MLOps, and autonomous ML agents.
28.3%
Career paths, role distinctions, portfolios, internships, hiring skills, and practical advice for becoming an analyst, data scientist, or data engineer.
21.7%
Statistical foundations and analytical methods, including regression, sampling, distributions, probability, experimental thinking, and inference.
21.7%
Visualization tools, dashboards, interactive reporting, and examples of communicating data through charts and real-time interfaces.
19.6%
Data pipelines, warehouses, orchestration, distributed systems, streaming, storage formats, and evolving data-stack tooling.
6.5%
Tone and stance
Performance benchmark
Posts with media make up 63% of this collection. Their median all-time score is 28.6, compared with 33.3 for text-only posts.
Format mix
Consensus and debate
Shared view
Several posts place problem definition, data understanding, and metric context before dashboards, modeling, or AI-generated recommendations. One AI-analysis post specifically recommends supplying definitions, changes, segments, objectives, history, and constraints as context.
Shared view
Learning posts prominently offer free textbooks, a 16-week analytics roadmap, coding-practice sites, and a regression explainer. Together, they emphasize statistics, technical fundamentals, and hands-on practice.
Shared view
Career-oriented posts recommend building applied projects, using projects as portfolio evidence, and tailoring dashboards to recurring problems in a target industry.
Open debate
Posts range from optimism about conversational and natural-language BI to explicit concerns about inaccurate or hallucinated AI analysis. One post reports a team reviewing AI analysis that was wrong “50% of the time”; another investment-focused post argues that novel AI findings merit extra scrutiny.
Open debate
One career narrative stresses adapting as stacks change through building projects, another prioritizes conceptual knowledge over tooling, and a role-to-tool map recommends selecting tools based on a target role.
What performs
The five supplied score outliers are learning-oriented posts: an analytics roadmap, free textbook list, coding-practice list, book recommendations, and a regression tutorial. Their all-time scores range from 434.06 to 1,526.04.
Lists are the most common supplied format, accounting for 20 of 46 tweets (43.5%), with a listed median all-time score of 89.69. The cited examples organize roadmaps, textbooks, and practice sites into scannable collections.
Stories account for 3 of 46 tweets (6.5%) and have a listed median all-time score of 94.979—higher than the listed medians for lists, tutorials, and opinions. The supplied examples cover a career path, AI-assisted report production, and a visualization-plugin build.
Statistical standouts
Creator landscape
The five most represented creators account for 21.7% of the selected posts.
1. Kirk Borne
@KirkDBorne
2 posts
2. Matt Dancho (Business Science)
@mdancho84
2 posts
3. Nick Singh | The Data Science Guy 📕
@NickSinghTech
2 posts
4. Towards Data Science
@TDataScience
2 posts
5. Zach Wilson
@Zachly
2 posts
6. Vaishnavi
@_vmlops
1 post
Matt Dancho shared a four-book data-science and AI reading list, plus business and leadership titles, and also posted a regression thread positioned for beginners.
Zach Wilson argues that conceptual knowledge matters more than tool names and uses a life-expectancy example to urge analysts to inspect underlying distributions rather than rely only on an average.
Kirk Borne shared a fraud-analytics guide covering descriptive, predictive, and social-network techniques, along with a resource on analytical skills for creating value from AI and data science.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 46-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Data Science tweets
Ranked 01–46
@Rita_tyna ·
LEARN DATA ANALYTICS FOR FREE! Here’s a structured 16-week roadmap for anyone who wants to learn data analytics but can’t afford paid courses right now. https://t.co/HjIy9QMuCD It contains resources on what data analytics is about, Maths/Basic Statistics, Excel, SQL, Power BI/Tableau, Python, AI for Data Analysts, How to create your Portfolio, LinkedIn and CV optimization for Data analysts as well as a 2 hour Data Analyst Interview Masterclass. Every resource included is free, detailed, and beginner-friendly.
@aiwithjainam ·
10 free textbooks from MIT, Stanford, and Berkeley that you can download legally right now. → Introduction to Linear Algebra - Gilbert Strang, MIT The textbook behind the most-watched math course in history. 20 million views on OCW. Every ML engineer learned this math from one quiet professor. https://t.co/Q5ZHXrBuD1 → Mathematics for Computer Science - MIT 6.042 Proofs, discrete math, probability. The actual foundation of CS that nobody tells undergrads about until it's too late. https://t.co/FOLUDXTubX → Convex Optimization - Stephen Boyd, Stanford Used in every serious ML and control systems course on earth. Cambridge University Press gave Boyd permission to keep it free on his own site. web. stanford. edu/~boyd/cvxbook/bv_cvxbook.pdf → CS229 Machine Learning Notes - Andrew Ng, Stanford Not the Coursera version. The actual Stanford graduate course notes. Dense, precise, and the closest thing to a grad school education you can download in one PDF. https://t.co/De8lcX59zt → An Introduction to Statistical Learning - Stanford / USC The book three statisticians from Stanford and USC made free because they wanted everyone to learn it. 290,000 people have taken the companion course on edX. https://t.co/TusQK9FDOc → Computational and Inferential Thinking - Berkeley Data 8 The textbook behind Berkeley's most popular course. Data science from scratch, built to be understood without a math degree first. https://t.co/xD2XAE49WA → Dive into Deep Learning - Berkeley / Amazon Jensen Huang called it "excellent." 500 universities across 70 countries use it. Every concept runs as live code directly in the browser. https://t.co/GJfBeDtJVO → Introduction to Probability - Blitzstein & Hwang, Harvard The official textbook of Harvard's Stat 110, which has been called the best probability course ever put on YouTube. Free second edition online. https://t.co/8dVzOUZDlH → The Elements of Statistical Learning - Hastie, Tibshirani, Friedman, Stanford The graduate-level version of ISLR. Springer makes it free as a PDF. Researchers keep a copy permanently in their downloads folder. https://t.co/PzdmjooaZW → MIT OCW Online Textbooks Index - 45+ books across every department One page. Every free MIT textbook organized by subject. Algorithms, physics, economics, engineering. All open access. https://t.co/eAjUcYzNa0 Save this before someone makes them take it down. (They won't. But save it anyway.)
@mdancho84 ·
You only need to read four books to truly get what’s going on in data science and AI: • Designing Machine Learning Systems by Chip Huyen • AI Engineering by Chip Huyen • Practical Statistics for Data Scientists by Peter Bruce, Andrew Bruce, and Peter Gedeck • Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow by Aurélien Géron If you read these four technical books and then read these four books on business value and leadership, you’ll be well on your way to career success: • The Lean Startup • Good Strategy Bad Strategy • The First 90 Days • The Hard Thing About Hard Things Any more you'd add?
@lennysan ·
Not enough people are talking about how much AI is impacting the role of data science. I was chatting with a DS friend, and he said that most of his team's work now is reviewing half-assed AI data analysis from PMs and engineers. And that 50% of the time, that analysis is wrong. The role is becoming less fun.
@heynavtoor ·
🚨 Google open sourced an AI that predicts the future. Stock prices. Sales trends. Energy demand. Weather patterns. Server traffic. Any time series. For free. It's called TimesFM. A foundation model built by Google Research specifically for time series forecasting. Published at ICML 2024. No training your own model. No data science degree. No expensive forecasting platforms. Feed it data. It predicts what happens next. Here's what makes this different from everything before it: → Pretrained on massive datasets. Works out of the box on YOUR data. → 200M parameters. Lightweight. Runs on a single GPU. → 16K context length. Feed it years of historical data in one pass. → Continuous quantile forecasting up to 1,000 steps ahead → Not just one prediction. Gives you confidence intervals. 10th to 90th percentile. → Works with PyTorch and JAX → Already deployed as an official Google product inside BigQuery Here's the wildest part: Traditional forecasting requires hiring data scientists, training custom models on your specific data, tuning hyperparameters for weeks, and praying it generalizes. TimesFM skips all of that. One pretrained model. Any domain. Any data. Just forecast. Give it stock prices. It predicts the trend. Give it server traffic. It predicts the spike. Give it sales data. It predicts the quarter. Give it energy demand. It predicts the grid. Bloomberg Terminal costs $25,000/year for forecasting tools. Enterprise forecasting platforms charge $50,000+ annually. Data science teams cost $500K+ in salary. pip install timesfm 10K GitHub stars. 825 forks. ICML 2024 paper. Apache 2.0 License. 100% Open Source. Built by Google Research.
@_vmlops ·
JPMORGAN OPEN-SOURCED THEIR INTERNAL PYTHON TRAINING used to train jpmorgan's own business analysts and traders now it's sitting on github with 13.2k stars, open for anyone ▫️ intro to numerical computing in python ▫️ data visualization with financial datasets ▫️ real financial data (iex cloud) + airport/route datasets ▫️ runs entirely in your browser via binder no setup needed ▫️ jupyter notebooks, taught by jpmorgan technologists no cs degree needed... built specifically for people without formal programming backgrounds if you're breaking into quant finance, data roles, or just want finance-focused python practice this is the repo https://t.co/EeFsuusvrw
@codyschneider ·
marketing engineering today how to get claude code to manage your facebook ads why it works - andromeda replaced interest targeting, the creative itself is now the targeting mechanism - ad and landing page are the signals facebook uses to decide who sees it - more ads across more personas gives the algorithm more to prospect with - loop: publish more ads, let facebook find winners, trim losers, promote winners, remix winners conversion tracking - google tag manager installed once in the head, all pixels deployed through it - custom event pushed to the data layer on the target user action - gtm sends that as a conversion event to facebook - start with an upfunnel conversion action, move deeper as volume allows - example progression: signup, then first product action, then payment - deeper events need enough volume for facebook to learn, that is why you cannot start at payment - cap at roughly five conversion events, more than that is noise server-side lead qualification - facebook conversions api fires events from a server - enrich the lead in real time, score 1 to 10, send back only leads above threshold (~6) - b2b signals: headcount, crm in use, site traffic (apify similarweb endpoint), branded search volume - smb signals: google review count and velocity - consumer signals: zip code - solves the lead quality complaint, the algorithm optimizes for form fills unless you tell it otherwise - avoids waiting on a four week sales cycle for signal - same approach works for google and linkedin research - scrape reddit for pain points and desired outcomes of the target buyer - use the exact language buyers already use in forums, do not guess - perplexity for manual pulls, xai api when the agent does it creative - statics: nano banana via kie ai (unified api across frontier image models) - ugc video: heygen avatars plus elevenlabs voices, elevenlabs voices outperform heygen native - seedance for higher quality, costs more - test with cheap models, produce winners with expensive ones - inputs: competitor ad library screenshots, your brand style guide (fonts, colors, logos) - ugc scripts written from the reddit research, hooks pulled from the same source - use heygen public avatars, ugc section marketing api - create a developer app in business manager, not a public app, no review process - create a system user, assign it the app, page, ad account, instagram account - generate a non-expiring token from the system user - use the api for writes only: uploading creative, turning off losers, moving ads between adsets - do not use the api for data analysis - hammering the api with agent read calls risks an ad account ban under their tos - agent reads also produce hallucination, truncation, and pagination errors at volume - mcp exposes only a handful of endpoints and breaks at scale, go direct to the api campaign structure - two campaigns: testing and winners - testing campaign holds multiple adsets, five ads per adset competing for budget - facebook concentrates budget on the likely winner on its own - signal window is two days to a week depending on spend and cost per action - winners move into the winners campaign and compete against each other there - watch for ad fatigue in the winners campaign - remix winners into new variations, push back into testing - optional: higher funnel event in testing, deeper event in winners analytics - data pipeline into a data warehouse, agent queries the warehouse - sources: facebook ads, google analytics or posthog, crm or cms, transaction data - open source stack: airbyte for the pipeline, clickhouse for the warehouse - postgres and supabase are not built for this, 100m rows takes minutes instead of milliseconds - facebook ads alone is about 42 tables with hundreds of columns - warehouse access enables conversational analytics and live dashboards deployment - an agent is software with a thinking loop and a live data stream - host on Graphed .com or similar, run hourly or daily
@techNmak ·
I've seen people spend $15,000 on AI/ML bootcamps and still not know this stuff. These playlists cover it for free. In the right order: 1./ Statistics & Data Analysis You can't model what you don't understand. Start here before you touch anything else. Playlist: https://t.co/rJkLQCRq3l 2./ Probability Bootcamp This is what your model is actually doing under the hood. Most people skip it. That's why most people stay stuck. Playlist: https://t.co/Y9GdEKwIA0 3./ Reinforcement Learning When you're ready to go beyond prediction into decision making. Intuitive. Not overwhelming. Playlist: https://t.co/rsL1s8ihyE 4./ Data Intensive Engineering Models are useless if they can't scale. This teaches you how to build systems that survive the real world. Playlist: https://t.co/0NNzeKEfsb Bookmark this. Thank yourself later.
@KanikaBK ·
MICROSOFT RESEARCH JUST PUT A FREE DATA ANALYSIS TOOL ONLINE THAT REPLACES $70/MONTH TABLEAU SEATS. You describe the chart you want. It builds it. No SQL or formulas. No degree required. 30 CHART TYPES. Works on screenshots, CSVs, live databases, and plain text. Zero dollars. Here is what is going on. Tableau charges $70 per user per month. Power BI Pro runs $10 to $20 per seat. Excel with Copilot is another subscription on top of that. Data analysis has always been expensive because the tools that make it easy cost serious money. Microsoft Research just put the alternative online for free. It is called Data Formulator. You load your data, describe what you want to see in plain language, and AI builds the chart, transforms the data, and writes the code behind it automatically. No dragging pivot tables. No writing SQL. No figuring out why your VLOOKUP broke. And here is where it gets interesting. ↳ paste a screenshot of a table and it extracts the data automatically ↳ connect it directly to MySQL, PostgreSQL, Azure, S3, or any URL with live refresh ↳ describe a chart in plain English and it figures out what data transformations are needed to make it ↳ agent mode lets it plan and explore your data across multiple turns on its own ↳ build full shareable reports directly inside the tool ↳ runs on OpenAI, Claude, Gemini, or fully local with Ollama The new version has a unified AI agent that handles everything, a persistent workspace so your data stays organized across sessions, and sandboxed code execution so nothing runs on your machine without permission. Microsoft sells Power BI to enterprises for real money every month. Their own research team then built a free open source tool that does a large chunk of the same job and put it on GitHub under an MIT license. Someone at the Power BI team is having a very interesting week.
@lemire ·
The Automation of Nonsense Among other things, I am the chair of the computer science programs at my university. Every year, I would write a report about what we had done and what we should do in the future. I wrote it in simple narrative prose, all by hand. Hardly anyone read my reports. Maybe nobody. But they were formally filed. Writing the narrative felt useful to me: it gave me a chance to reflect. I also got to present it for a few minutes, so I could share some insights. Yesterday, I learned that we are now required to follow a strict format for these reports, with multiple unintuitive sections and mandatory data analysis. The narrative is gone. I can no longer write a story. What used to be a somewhat enjoyable exercise has squarely entered the realm of “boring work you’d prefer not to do.” Of course, they provide nothing close to decent tools to generate the required data analysis. How anyone without a data science background is supposed to access the data, let alone analyze it properly, remains mysterious. A university is not a complicated place. You have students, programs, and courses. You can produce graphs and so forth. But what any of it actually means is hard to tell. So what did I do? I am a clever man, so I found a way to access data dumps of our historical data. I bet none of my colleagues know how to do that, but it is available if you know the URL and can navigate three submenus. I threw everything at an AI, including the description of the required analysis. I went into my archive, added all the files accumulated over the year, and simply prompted it. Then I waited… a few seconds… and there it was: a 42-page report complete with tables and everything. There were a few visible issues I had to fix, mostly because the AI lacked full context. For example, it noted that three courses had no enrollment records and correctly supposed that these courses had not yet been offered. All in all, it took me about two hours to produce a document that would have taken many weeks in the past. Is everything correct? It is too complicated to tell, but my colleagues skimmed it and expressed satisfaction. Are there mistakes? Probably. But had I done the work myself, given how ungrateful the task is, I likely would have made more mistakes. Importantly, the report serves no real purpose. We write it, some people skim it, we pass it around, and archive it. So why is it required? It is “proof of work.” I happen to be a good program chair, but many people do nothing at all. Forcing them to write a report about what they did creates pressure and gives the appearance of accountability. The obvious next step is to fully automate the process so that next year I just have to press a button. It defeats the “proof of work” angle. In my case, I did the work, so I am not concerned about providing the proof. What seems obvious is that my technique will be used by people who did not do the work. Do I feel bad about any of it? No, I do not. My new report is assuredly as good as or better than anything I could have done by hand. It goes much deeper into the data than I could have imagined. I even learned a few things. Many were intuitive beliefs, but it was satisfying to see the AI reach the same observations I would have made. We are losing something, of course: the genuine story told by a human being—what you are reading right now. But if you have worked in a bureaucracy, you know they don’t like that. They prefer cookie-cutter documents. The report is nonsense in the first place. Nobody needs these reports. In the words of Graeber, it is a bullshit job. And I just automated it.
@InduTripat82427 ·
What Have the World's Most Expensive Finance Teams Open-Sourced on GitHub? How Can Ordinary People Understand Quant Trading? Diving Right In Is the Fastest Way Top-tier quant and high-frequency trading firms like Jane Street, Goldman Sachs, J.P. Morgan, and others have released representative financial/engineering tools to help everyday quant enthusiasts learn institutional-grade pricing models, real-time data visualization, and high-precision performance debugging skills for free👇 1. Jane Street magic-trace (5.4k stars) https://t.co/bAgX7aqpKK A high-precision process tracing tool based on Intel Processor Trace. When ordinary profilers can't see the call stack clearly, it can record the complete execution process of every CPU instruction with nanosecond-level resolution. Strongly recommended for anyone wanting to dive deep into performance debugging and figure out exactly where the program is getting stuck 2. Goldman Sachs gs-quant (10.2k stars) https://t.co/C7wjgkVIXq A Python toolkit for derivatives pricing and risk management used daily by Goldman Sachs traders. It includes complete pricing models and risk calculation modules for common derivatives like options and swaps. You can install it directly with pip and start using it—perfect for those wanting to systematically learn institutional-grade quant pricing, with strong practical value 3. Perspective (originally a J.P. Morgan project, 10.5k stars) https://t.co/2NzmuQzZkZ J.P. Morgan's open-source powerhouse for real-time data visualization, especially adept at handling massive streaming market data. It lets you quickly build sleek interactive dashboards and real-time monitoring interfaces, supports Jupyter, and is more flexible than many paid terminals. Extremely friendly for folks doing data analysis and market visualization These three open-source projects let you directly access institutional-grade pricing models, real-time market dashboards, and high-precision performance debugging tools, helping ordinary developers boost their quant analysis, data visualization, and code optimization skills—all completely free
@Meer_AIIT ·
15 BEST GitHub Repos for AI&ML 1. Awesome Lists: https://t.co/G6douK0kyE 2. roadmap. sh: https://t.co/r52eb7oqUO 3. Python Data Science Handbook: https://t.co/A2C7OcxBpc 4. Machine Learning Notebooks, 3rd edition: https://t.co/Xqp3XH3eHp 5. Designing Machine Learning Systems (Chip Huyen 2022): https://t.co/CjIV5VuG0e 6. Neural Networks: Zero to Hero: https://t.co/6BSEwrAEhs 7. minGPT by karpathy: https://t.co/gyYTGi4Vnx 8. Project Based Learning: https://t.co/g9bli5FW2g 9. Build your own X: https://t.co/MW9OkM8GMa 10. awesome-generative-ai-guide: https://t.co/fmKKtocJXM 11. Made With ML: https://t.co/JQI0JSmG5J 12. Awesome Machine Learning: https://t.co/w43B9jXSLq 13. Awesome Data Science: https://t.co/qH0ULtqmAt 14. Awesome MLOps: https://t.co/wUs3bnxbxy h/t: yt Harry Connor AI
@parmardarshil07 ·
In 2018, I used CSV files and cron jobs. In 2019, I used SQL and Python. In 2020, I used Spark and AWS. In 2024, I used Airflow and Snowflake In 2026, I'm using AI agents to generate pipelines. 8 years. 8 completely different stacks. I wanted to become a data scientist. I dreamed about machine learning. I took every course — Andrew Ng, random Udemy ones, TensorFlow tutorials. I tried Kaggle problems and couldn't solve a single one. I'd open a dataset, stare at it, close my browser, and quit. So I took more courses. Thinking the NEXT one would finally unlock everything. It never did. Here's what actually changed things: ✅ I stopped consuming and started building. My first project was stupid simple — a classifier that detected exam notes in your camera roll and deleted them automatically. That led to an unpaid data science internship. Then, a data engineering internship I took just because I needed experience — any experience. I didn't even know what data engineering was. But here's what each phase actually taught me: → The web dev phase taught me how to write code that works in production → The course loop taught me that consumption without execution is a trap → The first project taught me that building > learning → The data science internship taught me NLP and how messy real problems are → The data engineering role taught me AWS, SQL, PySpark, and that I actually love the blend of code + business Here's the truth nobody tells you: Roadmaps don't work the way you think. Every data engineer I know has a completely different path. Mine started with PHP and ended up in Spark. Yours will look different too. What works: -> Learn Python and SQL (you'll use SQL 80-90% of the time) -> Learn big data fundamentals -> Build a project — even if it's ugly -> Apply for internships before you feel ready -> Share everything you learn in public ---- Your path won't be straight. Mine certainly wasn't. Tools expire. The ability to pick up the next one in a weekend? That's forever.
@Zachly ·
Conceptual knowledge is more important than tooling! Spark is a means of distributed compute Airflow is a means of job orchestration dbt is a means of data quality Tableau is a means of data visualization Iceberg is a means of data lake storage Flink is a means of stream processing Postgres is a means of consistent + available storage. MongoDB is a means of consistent + partition tolerant storage Cassandra is a means of available + partition tolerant storage Parquet is a means of data compression and serialization
@shub0414 ·
The only roadmap to go from 0 to ML/AI Expert Stage 1 – Python Basics Stage 2 – Statistics & Probability Stage 3 – Linear Algebra & Calculus Stage 4 – Data Preprocessing Stage 5 – Exploratory Data Analysis Stage 6 – Supervised Learning Stage 7 – Unsupervised Learning Stage 8 – Feature Engineering Stage 9 – Model Evaluation & Tuning Stage 10 – Deep Learning Basics Stage 11 – Neural Networks & CNNs Stage 12 – RNNs & LSTMs Stage 13 – NLP Fundamentals Stage 14 – Deployment (Flask, Docker) Stage 15 – Build projects Follow this roadmap to succeed!
@PythonDvz ·
Data Engineer vs Data Scientist: What’s the Difference? One builds the data foundation. The other turns data into intelligence. A Data Engineer designs pipelines, manages large-scale systems, ensures data reliability, and works heavily with cloud and distributed frameworks. They focus on performance, scalability, and architecture. A Data Scientist analyzes data, builds models, applies statistics, and translates patterns into actionable insights. They focus on prediction, experimentation, and business impact. If you enjoy system design, infrastructure, and data flow — engineering may suit you. If you enjoy analysis, modeling, and problem-solving with algorithms — science may be your path. Both roles are powerful. The real question is: do you want to build the engine or drive the strategy?
@goyalshaliniuk ·
Want to Learn Python for AI but Do not Know Where to Start? Here is a 20-step roadmap that takes you from complete beginner to building your first AI model in a structured, phase-by-phase journey. Whether you are aiming for data science, automation, or AI development, this roadmap gives you the exact sequence to master Python efficiently. PHASE 1: Python Fundamentals Lay the foundation by learning syntax, variables, loops, and functions. This phase helps you understand how Python works and prepares you for automation and AI tasks. PHASE 2: Data Structures & Libraries Discover how to organize, process, and visualize data using libraries like NumPy, Pandas, Matplotlib, and Seaborn — essential skills for AI development. PHASE 3: Data Preparation & Analysis Learn how to clean, explore, and transform data to make it AI-ready. This phase builds analytical thinking and introduces you to mini data projects. PHASE 4: Machine Learning Introduction Step into AI modeling with Scikit-learn. You’ll create regression and classification models, test their accuracy, and complete your first AI project from start to finish. Start Small, Stay Consistent You do not need years, just 20 focused steps. Follow this roadmap, code daily, and you’ll be ready to build your first AI model within weeks.
@joulee ·
A recent unlock for me on AI + data analysis: think less about prompting. Think more about cooking. See a lot of people use AI like a microwave. They drop in one chart, one problem statement, one KPI dip, and type: “Think like a senior analyst. What should I do?” Then they hit analyze and act surprised when what comes back is lukewarm slop. But good analysis is not microwave work. It’s chef work. If you give a great chef a microwave and say “make dinner,” you should not be shocked if the result is random. A chef needs more than that. They need a pantry. They need various tools. They need to know who they’re cooking for. They need to know whether this is Tuesday dinner or a wedding. They need to know what was already served. They need to taste as they go. They need constraints. Same with AI. Most people give AI one slice of the situation: “My growth is slowing. What should I do?” “Our retention is down. What’s happening?” “Revenue is up. Is that good?” That is not enough. Because a good answer depends on other context that narrows what is actually true. For example: What exactly is the metric? How is it defined? What changed recently? Which segments matter most? What are we optimizing for? What happened the last time this moved? What constraints are real? That’s what I mean by orthogonal context (which is a fancy way of saying, context that comes at right angles. That is independent from each other.) Different kinds of context that rule things in and out. This is why “better prompts” are overrated. “Act like a strategic analyst” is basically: “Cook like a Michelin chef.” The problem is not that the model is dumb. It’s that you gave it one thing and asked it to invent the meal. A better question is: What are the 5–7 things my best analyst would want to know before making a recommendation? Then, answer those questions. Give your AI the pantry and tools that it needs.
@heyitsalexP ·
I've been working to replace myself with AI since January. Here are my latest findings: Manus: best for big data analysis without hallucination and (obviously) Meta ads account analysis Also wonderful for synthesizing how business trends & Meta ads performance/changes tie together Needs context from you, or the outputs can resemble Meta account manager slop Lowest credit ceiling–seems like they're seeking profitability from the get-go vs other models Claude: Claude code is goated. Great for doing deep and comprehensive market and competitor research, then pivoting right into angle development and copy ideation. Still working to unlock the full potential of cowork. ChatGPT: Used to be my #1 rec for research, but performance has become less consistent and Claude is just as good (sometimes better) Fine for bouncing ideas around with no context, or generating prompts for images. No longer a "critical" tool for me. Parker: I am loving this tool more and more. I had to really put some structure behind my own creative strategy process to unlock the full potential. I use it for: quickly pulling ad performance metrics, brainstorming angle or persona ideation, asking it to "check my work" and evaluate concepts for creative diversity. As a competitive research tool, I'm still figuring out the best way to use it. Moby: great for generating weekly reports, but for ad hoc prompts it can be very slow. surprisingly good at concept ideation ("make me an ad that would appeal to men 45+, framing this product as a solution to slow metabolism") If I don't even know where to start defining a static format or sourcing competitors, sometimes Moby can do a decent job of creating something for a designer to work from.
@freeCodeCamp ·
The R programming language is a powerful tool for statistics and data analysis. And in this tutorial, Tiffany teaches you how to use R along with ggplot2 to create boxplots to model data. You'll inspect, clean, and prepare the data, perform exploratory data analysis, build some models, and more. https://t.co/TddLTySAQr
@petergyang ·
My top 5 takeaways from Sumeet (Brex) on building an AI data analyst with Claude Code: 1. Set up Claude Code to augment every step of data analysis Monitor dashboards and queries -> Explore metric changes -> Craft a good story -> Size potential impact. 2. The #1 mistake: Blowing up your context. To avoid this, use skills to enforce limits like "limit X rows on joins" and add timeouts that trigger query rewrites. 3. Build a skills + agents architecture for self-serve analysis Sumeet created skills for ad hoc analysis, data visualization, cohort analysis, and CSV exports (see below). Each skill enforces query limits (like “limit 50 on joins”), 2-3 minute timeouts, and standard patterns. This lets anyone run analyses without accidentally joining two million-row tables and crashing the database. 4. To build a great data analyst, you have to give Claude more context than just data. When metrics changed at Brex, Claude searched Slack, found an active incident, and connected it to the metric drop. “It saved me from asking what’s happening to our data?” 5. Cursor is crushing it for startups AND enterprise Brex’s data from Q4 shows Cursor consistently in the top 3 for spend. It’s both the startup and enterprise coding tool of choice. Would be curious to see if this changes in Q1 2026. 📌 Watch now: https://t.co/hvbtDwgjWM
@dbreunig ·
“taste” isn't enough… The three tiers of agent powered developers: 1️⃣ Can implement products: Can use agents to build code, with great tests. 2️⃣ Can implement products, 𝘄𝗶𝘁𝗵 𝗴𝗿𝗲𝗮𝘁 𝘁𝗮𝘀𝘁𝗲: Your feedback, which can keep up with the pace of code creation, lets you ship *good* products. 3️⃣ Has a deep empathy for their market and user and tons of surface area with them to stay up to date: Taste isn't enough! You need user empathy. Everyone keeps saying “taste”, to define a company's value when code is cheap. But “taste” is incomplete. User understanding, subject matter expertise, etc. These things are no longer the domain of just Product people, they need to be internalized by all and orgs should be designed to max out this surface area. I've said the difference between a good data scientist and a great data scientist is domain expertise: “When you have an idea how the business works, you can make more complex assumptions and develop hypotheses further out from the baseline. Bigger leaps, adequately tested, help you move faster and find unique information.” This now applies to software engineers.
@Zachly ·
My mentor Alex Hormozi made an inspiring quote that contains a data science error! He said “the average US males lives to 75. So you’re actually middle age at 37. So do what you need to do!” The problem with this stat is it assumes the US life expectancy is uniform when it’s clearly bimodal! When you do averages across a bimodal distribution, you get a number (75) that describes almost nobody in the US. America really is two countries. A nation with 3rd world life expectancy numbers (66-71 years) and one with Europe-level expectancies (81+ years) Avoid the red states on this chart and you’ll be fine making it to 80. Also whenever people make statistical claims, please consider the distribution of the underlying data!
@alliekmiller ·
We're seeing even more autonomous AI coworkers. The new MLE agent on the market is Disarray. In Kaggle competitions, Disarray: - won 28 medals across diverse domains (vision, NLP, tabular data) - placed top 10 in nine competitions - outperformed all human teams in one of those competitions ...each within 24 hours on a single GPU. The agent starts from a high-level task description and plans, runs, and refines ML workflows on its own and also grabs data beyond what it's given: it discovers and augments data using publicly available sources. Sam Altman recently predicted we would see an automated AI researcher in March 2028. And then you see stats like this and wonder if it will be earlier. Disarray backers include the co-founder of Databricks and Perplexity, the founder of Kaggle, the former U.S. Chief Data Scientist, and yours truly. Founders are two bad ass PhDs (ex-Databricks/Google/LinkedIn/MSFT, ex-NASA/IBM) that met at Cal.
@Al_Grigor ·
LLM systems feel like a new paradigm. In practice, much of the lifecycle still follows patterns that existed long before generative AI. One useful lens is CRISP-DM, a framework originally designed for data mining projects and widely adopted in data science. Even though the tools have changed, its phases map surprisingly well to how modern AI systems are built. Here is how the typical stages compare. 1. Business Understanding - Traditional ML: define the prediction task and success metrics. - AI systems: define the AI-powered product use case and the user experience you want to enable. 2. Data Understanding - Traditional ML: explore labeled datasets, distributions, and features. - AI systems: identify the inputs your system will use such as documents, images, APIs, databases, or external tools. 3. Data Preparation - Traditional ML: feature engineering, cleaning, and dataset curation. - AI systems: chunking documents, generating embeddings, building indexes, and wiring tools for agents. 4. Modeling - Traditional ML: train and tune models on structured datasets. - AI systems: prompt design, schema definition, retrieval pipelines, and agent behavior. 5. Evaluation - Traditional ML: metrics like accuracy, precision, and recall. - AI systems: task success, human feedback, and observable system behavior. 6. Deployment - Traditional ML: model serving and batch or online inference pipelines. - AI systems: full AI-powered applications that combine models, tools, and orchestration. The techniques look different, but the lifecycle remains largely the same. This is one reason many data scientists can smoothly transition into AI engineering roles. Read more about how CRISP-DM applies to AI Engineering: https://t.co/rkQw4fzVmT
@rohanpaul_ai ·
The power users of AI is pulling far ahead of the average employee. Workers in the 95th percentile of adoption generate 6X more AI messages than the median worker for basic chat tasks. The gap becomes much more extreme with advanced features. Among employees who work specifically in data analytics, these top-tier users interact with AI data analysis tools 16X more often than the median user in that same role. --- From 🌍 OpenAI's 2025 enterprise AI report openai .com/index/the-state-of-enterprise-ai-2025-report/
@predict_addict ·
The Statistical Test That Launched Spectral Analysis (1898) In 1898 — decades before Fisher, long before modern signal processing, and half a century before formal time-series theory — Arthur Schuster quietly solved a problem that still haunts data science: Read more: https://t.co/KBK6FNnwBQ
@Alex_TheAnalyst ·
Early in my career, I worked on double blind studies and we relied heavily on patient surveys - but patients don't always respond. So we would often sample data. Sampling is the practice of analyzing a subset of your data to draw conclusions about the whole. And it works surprisingly well when done right. Here's why sampling is useful: 1. Not everyone fills out a survey. Not every system logs every event. Sampling lets you work accurately with what you have instead of waiting on data that may never arrive. 2. A well-drawn sample accounts for gaps in your data so one missing group doesn't skew your entire analysis. 3. You don't need every data point to draw reliable conclusions. A properly selected sample can be just as accurate as the full dataset. Of course, there's a lot that goes into sampling correctly so you don't get incorrect results! I'm recording my full course on Statistics for Analyst Builder so statistics has been top of mind the past few weeks! I'll also be creating a Statistics Series on YouTube after the PostgreSQL series :)
@mukundiyngr ·
Who bet early on AI in oncology? Before “AI x Cancer” became consensus, a small group tied algorithms to clinical endpoints that could survive reality. Here’s what early conviction looked like inside NCI awards: ⬇️ - - - - - @theNCI | Peter Choyke (@PChoyke) 1ZIABC010655-18 Prostate Cancer Imaging 💡𝘈𝘐 𝘣𝘦𝘤𝘰𝘮𝘦𝘴 𝘳𝘦𝘢𝘭 𝘸𝘩𝘦𝘯 𝘪𝘵 𝘪𝘴 𝘵𝘪𝘦𝘥 𝘵𝘰 𝘪𝘮𝘢𝘨𝘪𝘯𝘨 𝘱𝘪𝘱𝘦𝘭𝘪𝘯𝘦𝘴 𝘢𝘯𝘥 𝘥𝘦𝘤𝘪𝘴𝘪𝘰𝘯-𝘮𝘢𝘬𝘪𝘯𝘨 @UCDavisHealth | Diana Miglioretti, PhD (@dmiglio ) 2P01CA154292-11 Advancing Equitable Risk-based Breast Cancer Screening 💡𝘳𝘪𝘴𝘬 𝘮𝘰𝘥𝘦𝘭𝘪𝘯𝘨 𝘵𝘩𝘢𝘵 𝘢𝘤𝘵𝘶𝘢𝘭𝘭𝘺 𝘸𝘰𝘳𝘬𝘴 𝘮𝘶𝘴𝘵 𝘸𝘪𝘯 𝘰𝘯 𝘦𝘲𝘶𝘪𝘵𝘺 𝘢𝘯𝘥 𝘩𝘢𝘳𝘮𝘴, 𝘯𝘰𝘵 𝘫𝘶𝘴𝘵 𝘈𝘜𝘊. @UTSWMedCenter | Steve Jiang (plus deep bench) 1R01CA276690-01A1 Spatial, genomic, pathologic biomarkers to guide ICI therapy (gastric cancer) 💡 𝘦𝘹𝘱𝘭𝘢𝘪𝘯𝘢𝘣𝘭𝘦 𝘮𝘶𝘭𝘵𝘪𝘮𝘰𝘥𝘢𝘭 𝘈𝘐 𝘮𝘰𝘷𝘦𝘴 “𝘱𝘳𝘦𝘥𝘪𝘤𝘵𝘪𝘰𝘯” --> “𝘵𝘳𝘦𝘢𝘵𝘮𝘦𝘯𝘵 𝘥𝘦𝘤𝘪𝘴𝘪𝘰𝘯𝘴.” @UNC | Susan Sumner @SumnerLab_NRI 5U24CA268153-02 Metabolomics and Clinical Assays Center 💡 𝘋𝘢𝘵𝘢 𝘪𝘴 𝘵𝘩𝘦 𝘮𝘰𝘢𝘵, 𝘢𝘯𝘥 𝘵𝘩𝘪𝘴 𝘪𝘴 𝘥𝘢𝘵𝘢𝘴𝘦𝘵 𝘪𝘯𝘧𝘳𝘢𝘴𝘵𝘳𝘶𝘤𝘵𝘶𝘳𝘦. @PennMedicine | Rinad Beidas & Robert Schnoll 5P50CA244690-02 Behavioral economics + implementation science to improve cancer care 💡 𝘢𝘥𝘰𝘱𝘵𝘪𝘰𝘯 𝘪𝘴 𝘢 𝘣𝘰𝘵𝘵𝘭𝘦𝘯𝘦𝘤𝘬. 𝘐𝘮𝘱𝘭𝘦𝘮𝘦𝘯𝘵𝘢𝘵𝘪𝘰𝘯 𝘴𝘤𝘪𝘦𝘯𝘤𝘦 𝘪𝘴 𝘢𝘯 𝘈𝘐 𝘮𝘶𝘭𝘵𝘪𝘱𝘭𝘪𝘦𝘳. @DanaFarber | Daphne Haas-Kogan (@DHaasKogan) 1U54CA274516-01A1 Pediatric radiation oncology at the interface of biology + data science 💡 multimodal cores + advanced dosimetry is how AI becomes a clinical capability @UCSFHospitals | Laura Esserman Grant: 1P01CA281826-01A1 WISDOM screening platform 💡 𝘴𝘤𝘳𝘦𝘦𝘯𝘪𝘯𝘨 𝘪𝘴 𝘸𝘩𝘦𝘳𝘦 𝘈𝘐 𝘵𝘩𝘢𝘵 𝘤𝘢𝘯 𝘴𝘩𝘪𝘧𝘵 𝘵𝘩𝘦 𝘤𝘶𝘳𝘷𝘦 𝘪𝘯 𝘴𝘤𝘳𝘦𝘦𝘯𝘪𝘯𝘨: 𝘸𝘩𝘰, 𝘩𝘰𝘸 𝘰𝘧𝘵𝘦𝘯, 𝘢𝘯𝘥 𝘸𝘪𝘵𝘩 𝘸𝘩𝘢𝘵 𝘮𝘰𝘥𝘢𝘭𝘪𝘵𝘺? @BoozAllen | Anna Fernandez (@AnnaFBiomed) 91022A00691022F00001-0-0-1 Low-dose CT images and corresponding data 💡𝘧𝘦𝘸 𝘴𝘢𝘺 𝘵𝘩𝘪𝘴: 𝘥𝘢𝘵𝘢𝘴𝘦𝘵𝘴 𝘢𝘳𝘦 𝘯𝘢𝘵𝘪𝘰𝘯𝘢𝘭 𝘢𝘴𝘴𝘦𝘵𝘴. 𝘞𝘪𝘵𝘩𝘰𝘶𝘵 𝘵𝘩𝘦𝘮, “𝘈𝘐 𝘪𝘯 𝘰𝘯𝘤𝘰” 𝘪𝘴 𝘵𝘩𝘦𝘢𝘵𝘦𝘳. vgbio (@PhysIQ) | Karen Larimer PhD, ACNP-BC, FAHA, FPCNA Larimer 75N91020C00040-P00001-9999-1 Wearable biosensors + AI analytics for early detection of decompensation 💡 𝘤𝘢𝘳𝘦 𝘪𝘴𝘯'𝘵 𝘫𝘶𝘴𝘵 𝘪𝘮𝘢𝘨𝘪𝘯𝘨 & 𝘱𝘢𝘵𝘩𝘰𝘭𝘰𝘨𝘺. 𝘊𝘰𝘯𝘵𝘪𝘯𝘶𝘰𝘶𝘴 𝘴𝘪𝘨𝘯𝘢𝘭𝘴 & 𝘦𝘢𝘳𝘭𝘺 𝘸𝘢𝘳𝘯𝘪𝘯𝘨 𝘴𝘺𝘴𝘵𝘦𝘮𝘴 𝘢𝘳𝘦 𝘤𝘰𝘮𝘪𝘯𝘨. @WUSTLnews | Daniel Marcus Grant: 5U24CA258483-03 Sustaining the I3CR imaging informatics center 💡 𝘪𝘯𝘧𝘰𝘳𝘮𝘢𝘵𝘪𝘤𝘴 𝘤𝘦𝘯𝘵𝘦𝘳𝘴 𝘢𝘳𝘦 𝘸𝘩𝘢𝘵 𝘮𝘢𝘬𝘦 𝘮𝘶𝘭𝘵𝘪-𝘴𝘪𝘵𝘦 𝘈𝘐 𝘳𝘦𝘱𝘳𝘰𝘥𝘶𝘤𝘪𝘣𝘭𝘦. - - - - - @OncoAlert Plot Methodology in fist comment below ⬇️
@onu_slim ·
Data Analysis Skills That Companies in Nigeria and Abroad Are Hiring For Data analysis has become one of the most reliable entry points into tech for beginners in Nigeria. Companies need people who can clean messy information, find useful patterns and present clear reports that support decisions. You can learn the core skills in three to six months and start earning. The foundation starts with strong Excel or Google Sheets skills. You must be comfortable with formulas, pivot tables, charts and basic data cleaning. From there, many people add Power BI or Tableau for creating professional dashboards. Learning basic SQL helps you pull data from databases, and a small amount of Python (using libraries like Pandas) opens more advanced opportunities. In Nigeria, junior data analysts and reporting officers commonly earn between N250,000 and N550,000 per month in full-time roles. Remote and freelance opportunities often pay higher. A single freelance dashboard or analysis project can range from N80,000 to N250,000 depending on complexity. Once you have a few successful projects and testimonials, monthly freelance income of N300,000 to N600,000 becomes realistic for consistent workers. To move from learning to earning, build three to five sample projects that solve real problems. Examples include sales performance dashboards, customer behaviour reports, or inventory analysis for a small business. Share these projects on LinkedIn and Twitter, and offer your service to local businesses or online clients at a beginner rate. Many people land their first paid work within four to six months of focused practice. The demand exists both inside Nigeria and with international clients who are comfortable working with remote talent. Start with one tool, master the basics, create visible proof of your ability, and begin offering the skill. Data analysis rewards clear thinking and consistency more than advanced theory. The companies are already looking for people who can turn numbers into useful insight.
@ethanrkho ·
Everyone's excited about AI in investing. Here's the paradox nobody talks about: Matei Zatreanu, founder of System2 (data science arm for $10B+ fundamental hedge funds, ex-King Street Capital) explains: "A fund manager got a perfect-looking AI response on market share. Every company had exactly equal share. The AI said, 'I didn't have any data, so I just made them all equal.'" "To generate alpha, you need to be right when everybody else is wrong. You're looking for outliers — something in the data everyone else is missing." "Here's the problem: when you need the data the most is when you're noticing an outlier. That's also when it's least reliable." "If the AI tells you something interesting, something novel — that's exactly when you should have the most concern about whether it's real or a hallucination." "We don't have infinite time or money. At some point you have to call it. That's why AI in fundamental investing is really, really hard." What part of the investment process should these tools be applied?
@iamKierraD ·
Become a data/business analyst. SQL, data visualization tool, excel. Build dashboards in the industry you want to be a data analyst in…focus them on common pain points in the industry. Ex: checking employee retention in HR, bed occupancy in hospitals(healthcare), etc Put this on resume and post about it.
@petesoder ·
How it feels to use new gen AI BI tools on company data. I’ve been drinking the @motherduck MCP and @_hex_tech Threads kool-aid and I’m starting to think we’ve turned a big corner for democratized data analysis. For years we’ve trained anyone who doesn't write SQL to think of data as an arms-length interaction with a dashboard or a sprint in an excel spreadsheet. In both cases there was a dead-end - some limitation that paused the curiosity loop. For most people in a company, the solution was to file a ticket and wait on a data eng/analyst but by then momentum is gone and some new task has taken priority. If insight is oil, you never drill deep enough because the loop is too long to keep drilling. Early “talk to your data” tools didn’t really solve this either. A lot of them were basically SQL generators dressed up in chat. Fine for trivia. Bad at accuracy. And not great for long threads of follow-ups. What feels different now is the emergence of BI tools that actually sustain their iterative curiosity loop, allowing people to deeper, quickly, without the painful task of recreating context each time. I think it's a big deal. It changes who engages with data, how often they do it and how much latent curiosity actually makes it into the system. And it’s only going to get better as new tools start to organize company context in a useful and portable way (i.e. Hex context studio, https://t.co/bp2d1L0U9T, Glean, Collate, et al). @barrald captures this shift well in the clip below. Worth a watch: https://t.co/GJ4JqZWHFt
@khalilApriday ·
𝗗𝗮𝘁𝗮 𝗥𝗼𝗹𝗲𝘀 vs 𝗧𝗼𝗼𝗹𝘀 — 𝗪𝗵𝗮𝘁 𝘁𝗼 𝗟𝗲𝗮𝗿𝗻 & 𝗪𝗵𝘆 One common mistake learners make 👇 Learning tools randomly without understanding the role they’re meant for. Here’s a quick, practical mapping of data roles to the tools they actually use: 🔹 Data Analyst → Excel, SQL, Power BI/Tableau, Pandas 🔹 Data Scientist → Python, SQL, Scikit-learn, Jupyter 🔹 ML Engineer → PyTorch/TensorFlow, Docker, Kubernetes, MLflow 🔹 Data Engineer → SQL, Spark, Kafka, Airflow, Cloud 🔹 AI Engineer → PyTorch, Hugging Face, APIs, Deployment tools 🔹 Business Analyst → Excel, BI tools, SQL, Presentations 🔹 Statistician → R/Python, StatsModels, SAS/SPSS 🔹 Data Architect → Cloud, Data Warehouses, Modeling tools 🔹 Research Scientist (AI/ML) → PyTorch/JAX, Colab, Experiment tracking 🔹 Big Data Engineer → Hadoop, Spark, Kafka, Databricks Key takeaway: 🎯 Don’t collect tools. 🎯 Pick a role → master the tools that role actually uses. Clarity in roles beats confusion in tools every time.
@NickSinghTech ·
Get more Tinder dates using Data Science 😈 This is a FANTASTIC portfolio project because: ✔️ they built interesting data visualizations ✔️ they solved their own pain point ✔️ they scraped + cleaned real-world data ✔️ they deployed their ML model to production
Best Tweets by Topic