<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:cc="http://cyber.law.harvard.edu/rss/creativeCommonsRssModule.html">
    <channel>
        <title><![CDATA[Stories by Kan Nishida on Medium]]></title>
        <description><![CDATA[Stories by Kan Nishida on Medium]]></description>
        <link>https://medium.com/@kanaugust?source=rss-1bfa80768afa------2</link>
        <image>
            <url>https://cdn-images-1.medium.com/fit/c/150/150/1*3EPEaiFps9Ds3GpqV8pNmA.png</url>
            <title>Stories by Kan Nishida on Medium</title>
            <link>https://medium.com/@kanaugust?source=rss-1bfa80768afa------2</link>
        </image>
        <generator>Medium</generator>
        <lastBuildDate>Fri, 09 Oct 2026 07:13:40 GMT</lastBuildDate>
        <atom:link href="https://medium.com/@kanaugust/feed" rel="self" type="application/rss+xml"/>
        <webMaster><![CDATA[yourfriends@medium.com]]></webMaster>
        <atom:link href="http://medium.superfeedr.com" rel="hub"/>
        <item>
            <title><![CDATA[How to Find the Right Number of Clusters with Silhouette Method]]></title>
            <link>https://medium.com/learn-dplyr/how-to-find-the-right-number-of-clusters-with-silhouette-method-1fe740064d45?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/1fe740064d45</guid>
            <category><![CDATA[data-analysis]]></category>
            <category><![CDATA[clustering]]></category>
            <category><![CDATA[data-science]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Thu, 25 Jun 2026 17:32:57 GMT</pubDate>
            <atom:updated>2026-06-25T17:32:57.072Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*SxrfnKJ-v9b6EoyW.png" /></figure><p>Clustering is one of the most useful techniques in exploratory data analysis.</p><p>Even when you do not know in advance what kinds of groups exist in your data, clustering can help you discover meaningful patterns such as:</p><ul><li>Customer segments</li><li>Product groups</li><li>Behavioral patterns</li><li>Regional differences</li><li>Types of survey respondents</li></ul><p>However, clustering always comes with one important question:</p><blockquote>How many clusters should we create?</blockquote><p>For example, in K-Means Clustering, one of the most widely used clustering algorithms, you need to decide the number of clusters, usually called K, before running the algorithm.</p><p>If K is too small, groups that should be separated may be combined into one cluster. If K is too large, groups that should belong together may be split into unnecessarily small pieces, making the result harder to interpret.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*ZsO2aAa6G2bHABCH.png" /></figure><p>This is why methods such as the Silhouette Method and the Elbow Method are often used to help choose a reasonable number of clusters.</p><p>In this article, we will focus on the Silhouette Method and discuss:</p><ul><li>What the Silhouette Method is</li><li>How to interpret the Silhouette Score</li><li>What to watch out for when using it</li><li>How it differs from the Elbow Method</li></ul><p>Exploratory recently added support for the Silhouette Method in K-Means clustering, so we will also look at how you can use it inside Exploratory at the end.</p><h3>What Is the Silhouette Method?</h3><p>The Silhouette Method is a way to evaluate how well each data point fits within the cluster it has been assigned to.</p><p>It answers two questions at the same time:</p><ol><li>Is this data point close to other data points in the same cluster?</li><li>Is this data point far away from data points in other clusters?</li></ol><p>A good clustering result should satisfy both conditions.</p><p>In other words, a data point should be:</p><ul><li>Similar to other points in the same cluster</li><li>Clearly different from points in other clusters</li></ul><p>A metric that measures how well it satisfy these two is called Silhouette Score.</p><h4>What Is the Silhouette Score?</h4><p>The Silhouette Score ranges from -1 to 1.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*hZ8FbsyLqs5sCO4v.png" /></figure><p>Let’s say we have several data points in a two-dimensional space like the one below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*AuN2HlzolASAqi5g.png" /></figure><p>Now, suppose we divide these data points into three clusters.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*jX23v3vEk0D7B_tw.png" /></figure><p>Let’s calculate the Silhouette Score for the point indicated by the arrow.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*_j144Xb3NIcy15_j.png" /></figure><p>The Silhouette Score can be calculated as follows.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*B73p2xi879NDudLc.png" /></figure><p>Cohesion measures how similar a point is to the other points in the same cluster. More specifically, it is calculated by summing the distances from that point to all the other points in the same cluster.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*KIx7cdwYnavBZA7a.png" /></figure><p>Separation, on the other hand, is the distance from that point to all the points in the nearest neighboring cluster. The larger this value is, the farther the point is from the other cluster.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*KxFCZDvqb5vxv8h4.png" /></figure><p>The Silhouette Score is calculated by subtracting cohesion from separation, then dividing the result by the larger of the two values.</p><p>For example, suppose the cohesion is 40 and the separation is 80.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*fHbXDRCFGcfKcBL3.png" /></figure><p>In this case, the Silhouette Score is 0.5.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*HYXAkTcqm1nXUnyI.png" /></figure><p>As you can see from the formula, when separation is large and cohesion is small, the numerator becomes larger, which makes the Silhouette Score higher. This means the point is closer to other points in its own cluster and farther away from points in other clusters.</p><p>On the other hand, when separation is small and cohesion is large, the numerator gets closer to zero or even becomes negative, which makes the Silhouette Score lower. This means the point is farther from other points in its own cluster and closer to points in another cluster.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*2NQANf39a4aOBRGz.png" /></figure><p>Hence the Silhouette Score ranges between -1 and 1, and -1 being the worst and 1 being the best.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*8bm7Km6NpzqbaG2d.png" /></figure><h4>Close to 1: Good</h4><p>A score close to 1 means the point fits well within its assigned cluster and is clearly separated from other clusters.</p><p>This is the ideal situation.</p><h4>Around 0: Ambiguous</h4><p>A score around 0 means the point is near the boundary between two clusters.</p><p>This does not always mean something is wrong, but it suggests that the separation between clusters is weak.</p><h4>Negative: Possible Misclassification</h4><p>A negative score means the point may be closer to another cluster than to the cluster it was assigned to.</p><p>If many data points have negative scores, you should be careful when interpreting the clustering result.</p><h3>Three Silhouette Metrics That Help Evaluate Clustering Quality</h3><p>So far, we have talked about the Silhouette Score for a single data point.</p><p>In practice, we calculate the score for all data points and then summarize the result. The most common summary is the average Silhouette Score.</p><p>However, looking only at the average can be misleading.</p><p>To evaluate clustering results more carefully, it is useful to look at three metrics together:</p><ul><li>Average Silhouette Score</li><li>Minimum Silhouette Score</li><li>Percentage of Negative Silhouette Scores</li></ul><p>Together, these metrics help you understand both the overall quality and whether there are problematic assignments.</p><h4>1. Average Silhouette Score: Overall Quality</h4><p>The average Silhouette Score is the average of the Silhouette Scores across all observations.</p><p>It answers the question:</p><blockquote>Overall, how well are the data points assigned to their clusters?</blockquote><p>For example, suppose we try different values of K and get the following results.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*dOfduZNVgWn06iE8gSzKGQ.png" /></figure><p>In this case, K = 3 has the highest average Silhouette Score, so it becomes a strong candidate.</p><p>However, the average can hide problems.</p><p>For example, some data points may be poorly assigned even if the overall average looks good. This is why we should also check the other two metrics.</p><h3>2. Minimum Silhouette Score: The Worst Assignment</h3><p>The Minimum Silhouette Score is the lowest score among all observations.</p><p>It answers the question:</p><blockquote>How bad is the worst cluster assignment?</blockquote><p>For example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ckGKqX3Ng-twQ89cy-Jn-g.png" /></figure><p>The average scores are almost the same, but K = 4 has a much worse minimum score.</p><p>This may indicate:</p><ul><li>An outlier</li><li>Noise in the data</li><li>Overlapping clusters</li><li>An unsuitable value of K</li></ul><p>That said, the minimum score is affected by a single observation. A very low minimum score does not automatically mean that the number of clusters is wrong.</p><p>It is best to treat it as a warning signal rather than a final decision rule.</p><h4>3. Percentage of Negative Silhouette Scores</h4><p>The percentage of negative Silhouette Scores tells you what percentage of observations have a score below 0.</p><p>It answers the question:</p><p>How many data points may be closer to another cluster than to their assigned cluster?</p><p>For example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WVe0R3xffjb4Y9HbZPFkLA.png" /></figure><p>In this case, K = 4 has a slightly higher average score, but it also has a much higher percentage of negative scores.</p><p>That is a warning.</p><p>It may mean that K = 4 creates clusters that look good on average but leave many observations poorly assigned.</p><p>In this situation, K = 3 may be a more stable and reliable choice, even though its average score is slightly lower.</p><h4>How to Use the Three Metrics Together</h4><p>Here is a simple way to think about the three metrics.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*YIHgG-22fX2N8vVMo5ekOg.png" /></figure><p>A practical workflow is:</p><ol><li>Start with the K that has the highest average Silhouette Score.</li><li>Check whether the percentage of negative scores is low.</li><li>Check whether the minimum score is not extremely low.</li><li>Compare the actual cluster profiles.</li><li>Choose the K that is meaningful and useful for your analysis.</li></ol><p>The goal is not simply to choose the highest score, it is to find clusters that balance among:</p><ul><li>Separation</li><li>Stability</li><li>Interpretability</li><li>Practical usefulness</li></ul><h3>Silhouette Method vs. Elbow Method</h3><p>Another popular method for choosing the number of clusters is the Elbow Method.</p><p>Silhouette Method and Elbow Method are both useful, but they evaluate different aspects of clustering.</p><h4>Elbow Method</h4><p>The Elbow Method looks at how compact the clusters are.</p><p>In K-Means clustering, the algorithm tries to minimize the distance between each data point and its cluster center. This distance is often summarized as within-cluster sum of squares, or WSS.</p><p>As K increases, WSS always decreases.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*DYiJXBdD2QUSn3xj.png" /></figure><p>As we increase the number of clusters from 1 to 2, then from 2 to 3, the distance (WSS) tends to decrease substantially at first. However, after a certain point, the amount of decrease starts to level off.</p><p>The point where this decrease changes from steep to gradual is called the “elbow,” and the K value, or number of clusters, at this elbow is considered a good candidate for the optimal number of clusters.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*KWT0kZsVP47MAZMA.png" /></figure><p>The challenge is that, in real-world data, the elbow is not always clear. Sometimes the curve gradually decreases without an obvious bend.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*6zqfO9Utk2TGOouX.png" /></figure><h4>Silhouette Method</h4><p>The Silhouette Method, on the other hand, evaluates how well clusters are separated.</p><p>It considers both:</p><ul><li>How close points are within the same cluster</li><li>How far they are from neighboring clusters</li></ul><p>In simple terms:</p><ul><li>The Elbow Method asks, “How compact are the clusters?”</li><li>The Silhouette Method asks, “How well separated are the clusters?”</li></ul><p>Both are useful, and in practice it is often best to look at both.</p><p>However, when comparing different values of K, the Silhouette Method has a practical advantage because its score is easier to interpret and compare.</p><p>This is especially true when you look at not only the average score, but also the minimum score and the percentage of negative scores.</p><h3>A Practical Workflow for Choosing the Number of Clusters</h3><p>Here is a practical way to use the Silhouette Method.</p><h4>1. Choose a Range of K Values</h4><p>Start with a reasonable range, such as K = 2 to K = 10.</p><p>Technically, you can create many clusters, but if the result becomes too hard to explain, it may not be useful in practice.</p><h4>2. Compare the Average Silhouette Scores</h4><p>As a rough guideline:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*kN6tJkyMvtnjYyImIYxewQ.png" /></figure><p>However, these are only rough guidelines.</p><p>For survey data, behavioral data, and many real-world business datasets, scores can often be lower. A “weak” score does not automatically mean the clusters are useless.</p><p>The real question is whether the clusters are meaningful for your purpose.</p><h4>3. Check the Percentage of Negative Scores</h4><p>Look at how many observations may be poorly assigned.</p><p>If one K has a slightly higher average score but a much higher percentage of negative scores, you may want to be cautious.</p><h4>4. Check the Minimum Score</h4><p>The minimum score can help you identify outliers or problematic assignments.</p><p>Again, do not overreact to a single low value, but use it as a warning signal.</p><h4>5. Compare Similar Candidate K Values</h4><p>Do not automatically choose the K with the highest average Silhouette Score.</p><p>If several values of K have similar scores, compare them as candidates.</p><h4>6. Profile the Clusters</h4><p>Once you have a few candidate K values, create the clusters and examine their characteristics.</p><p>For example, in customer segmentation, you may want clusters that can be explained in simple language, such as:</p><p>“High-frequency customers with high spending and low discount usage.”</p><p>If the clusters are hard to explain, they may not be useful even if the score looks good.</p><h4>7. Use Domain Knowledge</h4><p>Clustering is an unsupervised learning method, which means there is no correct answer label.</p><p>Unlike regression or classification, we cannot simply compare predictions against known outcomes.</p><p>This means the final decision should be made by combining quantitative metrics, visual analysis, and domain knowledge.</p><p>The Silhouette Method supports your decision, but it does not replace your judgment.</p><h3>Using the Silhouette Method in Exploratory</h3><p>Starting with Exploratory v15.5, you can use the Silhouette Method in the K-Means Clustering under Analytics view.</p><h4>Choosing the most optimized number of clusters</h4><p>When you run K-Means clustering, Exploratory automatically evaluates multiple values of K using the Silhouette Method and show you the results.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*k2eevVTuuF06n2h3.png" /></figure><p>You can compare different cluster counts and choose reasonable candidates.</p><p>For example, suppose you analyze this employee survey data, which you can download from <a href="https://exploratory.io/data/kanaugust/spu8ZuT1pR">here</a>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*7DBS_sfF-1HdXr7q.png" /></figure><p>The average Silhouette Score may suggest that K = 2 is the best option, while K = 3 is the second-best option. However, when you look at the percentage of negative scores, you may find that K = 2 has more observations that are not clearly separated.</p><p>In this case, you may want to compare K = 2 and K = 3 by looking at the actual cluster profiles.</p><p>In Exploratory, you can use visualizations such as radar charts, box plots, and scatter plots to compare the characteristics of each cluster.</p><p>For example, with K = 2, the result may simply separate respondents into a “high satisfaction” group and a “low satisfaction” group.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*OgqtKrzT0RcOVttc.png" /></figure><p>On the other hand, with K = 3, the result may further separate the high-satisfaction respondents into two different patterns.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*ckd_-swcs8H1G40D.png" /></figure><p>In this case, K = 3 may be more useful, even if K = 2 has a slightly better average score.</p><p>This is where cluster interpretation becomes important.</p><p>You can also use Exploratory’s AI Summary to get help interpreting the clustering result, including advice on the number of clusters and explanations of the cluster characteristics.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*oprGl-nHMNaCl5Vk.png" /></figure><h4>Checking Silhouette Scores by Cluster</h4><p>Exploratory also shows Silhouette Score metrics for each cluster.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*uGqBFICHSM8qnOdv.png" /></figure><p>This is useful because the overall score may hide differences across clusters.</p><p>For example, some clusters may be well separated, while another cluster may contain many ambiguous observations.</p><p>By checking the average Silhouette Score and the percentage of negative scores for each cluster, you can identify which clusters are stable and which ones should be interpreted with caution.</p><p>This helps you avoid treating all clusters as equally reliable.</p><h3>Summary</h3><p>Clustering is useful for many types of analysis, but choosing the number of clusters can be confusing.</p><p>The Silhouette Method helps you by answering questions such as:</p><ul><li>Are the clusters clearly separated?</li><li>Are we creating too many groups?</li><li>Are there many poorly assigned observations?</li><li>Which K gives us a more stable and interpretable result?</li></ul><p>The key is not to choose the highest score alone.</p><p>Instead, use the Silhouette Method as part of a broader workflow that combines:</p><ul><li>Quantitative evaluation</li><li>Visual exploration</li><li>Cluster profiling</li><li>AI-assisted interpretation</li><li>Human judgment</li></ul><p>What I personally find exciting about the new Silhouette Method support in Exploratory is that it connects these steps into one workflow.</p><p>You can evaluate the clustering result statistically, explore the clusters visually, get interpretation support from AI, and then make the final decision based on your own knowledge and purpose.</p><p>That is the kind of workflow that helps people use data to make better decisions.</p><h3>Try with Exploratory</h3><p>If you have wanted to try cluster analysis, or if you have used clustering before but struggled with choosing the right number of clusters, this is a good time to try it in Exploratory.</p><p>👉 Download the latest version of Exploratory</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you do not have an Exploratory account yet, you can sign up and start a 30-day free trial here:</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>If your trial has already expired, you can launch the latest version and click “Extend Trial” to try it again.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=1fe740064d45" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/how-to-find-the-right-number-of-clusters-with-silhouette-method-1fe740064d45">How to Find the Right Number of Clusters with Silhouette Method</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Introduction to CatBoost: Why It Works So Well with Categorical Data]]></title>
            <link>https://medium.com/learn-dplyr/introduction-to-catboost-why-it-works-so-well-with-categorical-data-96c41dc2cd9e?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/96c41dc2cd9e</guid>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[machine-learning]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Sat, 13 Jun 2026 07:36:20 GMT</pubDate>
            <atom:updated>2026-06-13T07:36:20.430Z</atom:updated>
            <content:encoded><![CDATA[<h4>A practical guide to how CatBoost handles categorical variables, how it differs from XGBoost and LightGBM, and when to use it for real-world business data.</h4><p>In recent years, gradient boosting algorithms such as XGBoost, LightGBM, and CatBoost have become popular choices for building accurate predictive models on tabular business data.</p><p>At Exploratory, we added support for CatBoost in version 15.5, based on requests from our users.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*K7sKZLz8tP7k5rrP.png" /></figure><p>So, what is CatBoost?</p><p>To understand why CatBoost is useful, it helps to look at the kind of data we often work with in the real world.</p><p>Business data does not always come neatly packaged as numbers.</p><p>Customer data often includes fields such as customer segment, region, industry, and subscription plan. Marketing data may include campaigns, ad channels, and traffic sources. Survey data often contains response categories and demographic attributes.</p><p>In other words, business data is full of categorical variables.</p><p>CatBoost is a machine learning algorithm designed to build high-performing predictive models, especially when your data contains many categorical variables.</p><p>In this post, I’ll explain what CatBoost is, how it differs from XGBoost and LightGBM, and why it is especially useful for business data with many categories.</p><p>CatBoost can be particularly effective when your data includes variables such as:</p><ul><li>Customer segment</li><li>Region</li><li>Product category</li><li>Industry</li><li>Subscription plan</li><li>Campaign</li><li>Ad channel</li><li>Survey responses</li></ul><p>We’ll cover:</p><ul><li>What problem CatBoost was designed to solve</li><li>How it differs from Random Forest, XGBoost, and LightGBM</li><li>Why CatBoost handles categorical variables well</li><li>What kinds of data CatBoost is good for</li><li>When you should consider using CatBoost</li></ul><p>If you’d like to learn more about LightGBM, you may also find this article helpful:</p><ul><li><a href="https://exploratory.io/note/kanaugust/XSg9wsR6eK">Introduction to LightGBM: How It Differs from Random Forest and XGBoost</a></li></ul><h3>The Problem CatBoost Tries to Solve</h3><p>Suppose you want to predict customer churn based on data like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CPRn3Kl6tI5nJavgvOMhkw.png" /></figure><p>In this data, Monthly Revenue is numeric, but everything else is categorical.</p><p>In many real-world business datasets, categorical variables are not a minor. They are often a major part of the data.</p><p>But most machine learning algorithms cannot directly understand text values such as:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*wpf8_SpL1CS_7noS9oFFcQ.png" /></figure><p>These values need to be converted into numbers before they can be used by a model.</p><p>This is where CatBoost starts to become interesting.</p><p>Before going deep down into CatBoost, let’s refresh our memory on how Boosting model has evolved over the last decades, starting from Random Forest to CatBoost.</p><h3>Random Forest: Many Independent Trees</h3><p>Random Forest builds many decision trees independently and combines their results.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*RINQ2Qs4XuppgxIJ.png" /></figure><p>Each tree follows a process like this:</p><ol><li>Randomly sample rows from the data.</li><li>Randomly select a subset of variables when splitting.</li><li>Let each tree make its own prediction.</li></ol><p>The final prediction is based on the average prediction for regression, or the majority vote for classification.</p><p>The core idea is simple: by combining many independent trees, Random Forest reduces the instability of a single decision tree and produces more stable predictions.</p><p>However, the trees do not learn from each other.</p><p>Each tree is built independently. One tree does not try to fix the mistakes made by another tree.</p><p>As a result, Random Forest is often easy to use and fairly robust, but it may not achieve the same level of predictive accuracy as modern boosting algorithms.</p><h3>XGBoost: Trees That Correct Previous Mistakes</h3><p>XGBoost uses a method called gradient boosting.</p><p>Instead of building many trees independently, it builds trees sequentially. Each new tree tries to correct the mistakes made by the trees that came before it.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*5P7wgsHnzYVBvpFu.png" /></figure><p>For example, in a regression problem, the first tree makes an initial prediction. Then the model calculates the errors, or residuals, between the actual values and the predicted values.</p><p>RowActual ValuePredicted ValueResidual11073215141389–1</p><p>The next tree is trained not to predict the original target directly, but to predict these residuals.</p><p>Why?</p><p>Because the residuals tell the model how it should adjust its predictions to reduce the loss.</p><p>Conceptually, the model evolves like this:</p><ul><li>Tree 1 → Initial prediction</li><li>Tree 2 → Corrects the errors from Tree 1</li><li>Tree 3 → Corrects the remaining errors</li><li>Tree 4 → Continues improving the prediction</li></ul><p>The updated prediction can be thought of as:</p><blockquote>New prediction = Prediction from previous trees + Learning rate × Prediction from the new tree</blockquote><p>This process often produces higher accuracy than Random Forest.</p><p>However, as data becomes larger and more complex, building many trees sequentially can become computationally expensive.</p><h3>LightGBM: Built for Speed and Scale</h3><p>LightGBM is also a gradient boosting algorithm, but it was designed to make boosting faster and more scalable.</p><p>It uses several techniques to improve performance, including:</p><ul><li>Histogram-based splitting</li><li>Leaf-wise tree growth</li><li>Gradient-based One-Side Sampling, or GOSS</li><li>Exclusive Feature Bundling, or EFB</li></ul><p>LightGBM is especially useful when:</p><ul><li>The dataset has many rows.</li><li>There are many explanatory variables.</li><li>Training speed matters.</li><li>You need to scale to larger data.</li></ul><p>Because of this, LightGBM has become one of the most widely used machine learning algorithms for practical business applications.</p><p>For more details, see:</p><ul><li><a href="https://exploratory.io/note/kanaugust/XSg9wsR6eK">Introduction to LightGBM: How It Differs from Random Forest and XGBoost</a></li></ul><h3>So What Makes CatBoost Different?</h3><p>CatBoost is also a gradient boosting algorithm, just like XGBoost and LightGBM.</p><p>The overall learning process is similar: it builds many trees sequentially, and each new tree helps improve the predictions made so far.</p><p>But CatBoost has one major design focus:</p><blockquote><em>CatBoost was designed to handle categorical variables effectively.</em></blockquote><p>This makes it especially useful when your data includes categorical variables with many possible values, such as:</p><ul><li>Product categories</li><li>Campaign IDs</li><li>Regions</li><li>Stores</li><li>User segments</li><li>Features used by customers</li><li>Survey responses</li></ul><p>These are sometimes called high-cardinality categorical variables.</p><p>To understand why CatBoost matters, let’s first look at the common ways categorical variables are handled.</p><h3>The Problem with One-Hot Encoding</h3><p>A common way to handle categorical variables is one-hot encoding.</p><p>For example, suppose you have a subscription plan column like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_y0nGF5aZPwQWdoArmUKKw.png" /></figure><p>One-hot encoding turns each category into a separate column:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xmHHI_Nu0s_TQvN5v77eow.png" /></figure><p>This works fine when the number of categories is small.</p><p>But what if you have:</p><ul><li>500 product categories</li><li>1,000 regions</li><li>10,000 campaign IDs</li></ul><p>Then one-hot encoding creates a very large number of columns.</p><p>This can cause several problems:</p><ul><li>Higher memory usage</li><li>Slower training</li><li>Sparse data</li><li>Higher risk of overfitting</li></ul><p>So one-hot encoding is not always ideal, especially for business data with many categorical variables.</p><h3>Target Encoding: A More Informative Alternative</h3><p>Another approach is called target encoding.</p><p>Instead of turning each category into many binary columns, target encoding replaces each category with a statistic calculated from the target variable.</p><p>For example, if you are predicting churn, you might calculate the churn rate for each subscription plan:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ocCiR54BoRFN1deGfzt5Sw.png" /></figure><p>Then the model can use this information directly.</p><p>For example:</p><blockquote>Starter → 0.67<br>Business → 0.00</blockquote><p>This gives the model useful information:</p><p>“Customers on the Starter plan tend to churn more often.”</p><p>Target encoding can be very powerful, especially for categorical variables with many levels.</p><p>But it also has a serious risk.</p><h3>The Risk of Data Leakage</h3><p>Suppose you have this small dataset:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*793siy2lKrurkxWiPh9rIg.png" /></figure><p>If you calculate the churn rate for the Starter plan, you get:</p><blockquote>2 / 3 = 0.67</blockquote><p>So you might encode the Starter plan as:</p><blockquote>Starter → 0.67</blockquote><p>But there is a problem.</p><p>When you calculate the value 0.67, you are using the target value, Churn, from all rows, including the row you are trying to predict.</p><p>For Customer A, the encoded value includes Customer A’s own churn result.</p><p>That means you are using the answer to create an input variable.</p><p>This is a form of <strong>data leakage</strong>.</p><p>When this happens, the model may look very accurate on the training data, but it may perform poorly on new data because it has learned from information that would not be available in a real prediction setting.</p><p>This is one of the problems CatBoost was designed to address.</p><h3>Ordered Target Statistics</h3><p>CatBoost uses a technique called Ordered Target Statistics to reduce leakage when encoding categorical variables.</p><p>The idea is simple:</p><blockquote><em>When calculating the encoded value for a row, use only the rows that come before it.</em></blockquote><p>Let’s use the same example.</p><p>Suppose the data is arranged in a random order:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*P_N2bAnBpSkmb8SSB2Lflg.png" /></figure><p>Assume the overall churn rate, used as the prior, is 50%.</p><p>CatBoost converts the Plan variable like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*N-gdFYXhJdQKcZHWoW15sg.png" /></figure><p>What is happening here?</p><p>For Customer A, there are no previous Starter customers, so CatBoost uses the prior:</p><blockquote>A → No previous Starter customers → Use 0.50</blockquote><p>For Customer B, the only previous Starter customer is A:</p><blockquote>B → Previous Starter customers: A → Churn rate = 1 / 1 = 1.00</blockquote><p>For Customer C, the previous Starter customers are A and B:</p><blockquote>C → Previous Starter customers: A and B → Churn rate = (1 + 0) / 2 = 0.50</blockquote><p>For Customer D, the previous Starter customers are A, B, and C:</p><blockquote>D → Previous Starter customers: A, B, and C → Churn rate = (1 + 0 + 1) / 3 = 0.67</blockquote><p>The key point is this:</p><p>The encoded value for each row is calculated without using that row’s own target value.</p><p>So CatBoost can still use the useful information that Starter customers tend to churn more often, while reducing the leakage that can happen with naive target encoding.</p><p>This is one of the main reasons CatBoost works well with categorical variables.</p><h3>Ordered Boosting</h3><p>CatBoost also uses another important idea called Ordered Boosting.</p><p>In standard gradient boosting, the model is trained on the full training data. Then it makes predictions on that same training data and calculates the errors. The next tree is trained to correct those errors.</p><p>At first glance, this seems reasonable.</p><p>But there is a subtle issue.</p><p>When the model predicts a row in the training data, that row has already been used to train the model. In other words, the model is predicting a row it has already seen.</p><p>In production, the situation is different. The model is asked to predict new data that it has not seen before.</p><p>This difference between the training-time prediction situation and the production-time prediction situation is sometimes called <strong>prediction shift</strong>.</p><p>CatBoost tries to reduce this shift by using a random order of the data.</p><p>When calculating the prediction for a row, CatBoost uses only the data that comes before that row in the random order.</p><p>For example, if the data is ordered as A, B, C, D, then when calculating the error for C, the model uses only A and B. It does not use C’s own target value to help predict C.</p><p>This makes the training process more similar to the real-world prediction process, where the model must predict data it has not seen before.</p><p>Ordered Target Statistics helps reduce leakage when encoding categorical variables.</p><p>Ordered Boosting helps reduce prediction shift during boosting.</p><p>Together, these ideas are a big part of what makes CatBoost different from XGBoost and LightGBM.</p><h3>Symmetric Trees</h3><p>Another important feature of CatBoost is that it uses symmetric trees, also called oblivious trees.</p><p>In a regular decision tree, different branches can use different split conditions.</p><p>Conceptually, it may look like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*m7NCQf8iZ5sKPkTo.png" /></figure><p>In a symmetric tree, all nodes at the same depth use the same split condition.</p><p>Conceptually, it looks like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*YIvDDNLv8cyW9FmW.png" /></figure><p>This means, one of the trees would look something like this.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*dV5UbixgjiQZ5EKn.png" /></figure><p>At first, this may sound restrictive.</p><p>A regular decision tree can be more flexible because each branch can choose its own split condition.</p><p>But CatBoost is not trying to build one perfect tree.</p><p>It is a boosting model. It builds many relatively small trees and combines them.</p><p>One tree may capture a broad pattern. The next tree may correct part of the remaining error. Later trees continue refining the prediction.</p><p>So even if each individual symmetric tree is less flexible, the full boosted model can still be very expressive.</p><p>The symmetric tree structure also has several advantages:</p><ul><li>Faster prediction</li><li>Better memory efficiency</li><li>Simpler and more stable tree structure</li><li>Lower risk of overfitting</li><li>Good performance when many trees are combined</li></ul><p>This is another example of CatBoost’s design philosophy: use a more controlled tree structure, then gain flexibility by combining many trees through boosting.</p><h3>What CatBoost Cannot Protect You From</h3><p>At this point, it may sound like CatBoost solves data leakage.</p><p>But that is not quite true.</p><p>CatBoost helps reduce leakage caused by target encoding of categorical variables.</p><p>It does not protect you from leakage already present in the data itself.</p><p>For example, suppose you are predicting customer churn and you include variables such as:</p><ul><li>Churn date</li><li>Refund amount after cancellation</li><li>Number of support tickets after cancellation</li><li>Final billing status</li><li>Cancellation reason</li></ul><p>These variables would not be available at the time you actually need to make the prediction.</p><p>If you include them, any model can produce artificially high accuracy.</p><p>CatBoost cannot fix that.</p><p>The basic rule remains the same:</p><blockquote><em>Use only the information that would be available at the time of prediction.</em></blockquote><p>This is true for CatBoost, LightGBM, XGBoost, Random Forest, and any other predictive model.</p><h3>How XGBoost and LightGBM Handle Categorical Variables</h3><p>By the way, it is important to clarify one point here.</p><p>Using categorical variables with XGBoost or LightGBM does not automatically mean you have a leakage problem.</p><p>XGBoost, LightGBM, and CatBoost use different strategies for categorical variables.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_78J8ax1vmN48FTAwMGbvg.png" /></figure><p>For example, LightGBM can learn category-based splits such as:</p><pre>Plan in {Starter, Pro}</pre><p>versus:</p><pre>Plan in {Business, Enterprise}</pre><p>The resulting rule may look like:</p><pre>if Plan is Starter or Pro:<br>    go left<br>else:<br>    go right</pre><p>CatBoost takes a different approach.</p><p>Instead of only grouping categories, it can convert a category into a target-based statistic in a leakage-aware way.</p><p>For example:</p><pre>Starter<br>↓<br>Average churn rate among previous Starter customers</pre><p>The important point is that these models have different design goals.</p><p>LightGBM emphasizes speed and scalability.</p><p>CatBoost emphasizes effective use of categorical variables while reducing leakage and prediction shift.</p><p>XGBoost is a strong general-purpose boosting algorithm with broad adoption and flexibility.</p><h3>What Kind of Data Is CatBoost Good For?</h3><p>CatBoost is especially useful for business datasets that contain many categorical variables.</p><p>Examples include:</p><h4>Customer Data</h4><ul><li>Customer segment</li><li>Region</li><li>Industry</li><li>Subscription plan</li><li>Account type</li></ul><h4>Marketing Data</h4><ul><li>Ad channel</li><li>Campaign ID</li><li>Traffic source</li><li>Landing page</li><li>Creative type</li></ul><h4>Product Data</h4><ul><li>Product category</li><li>Brand</li><li>SKU</li><li>Store</li><li>Supplier</li></ul><h4>Survey Data</h4><ul><li>Demographic attributes</li><li>Response categories</li><li>Segments</li><li>Preference groups</li></ul><h4>SaaS Data</h4><ul><li>Plan type</li><li>Customer segment</li><li>Feature usage category</li><li>Acquisition channel</li><li>Support category</li></ul><p>If your dataset has many categorical variables, especially variables with many unique values, CatBoost is often worth trying.</p><h3>Which Model Should You Use?</h3><p>So which model should you choose: Random Forest, XGBoost, LightGBM, or CatBoost?</p><p>In practice, the best answer is usually:</p><blockquote><em>Try multiple models, evaluate them properly, tune the parameters, and choose the one that performs best on validation or test data.</em></blockquote><p>That said, the following guideline can be helpful as a starting point:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6NerxhKY-hyoOpkrH_r0-w.png" /></figure><p>This does not mean one model is always better than the others. It means each model has a different strength.</p><p>Random Forest is stable and easy to use.</p><p>XGBoost is powerful and flexible.</p><p>LightGBM is fast and scalable.</p><p>CatBoost is strong when categorical variables play an important role.</p><p>For many business datasets, especially customer, marketing, survey, and SaaS data, CatBoost is a very practical option to consider.</p><h3>Summary</h3><p>Random Forest, XGBoost, LightGBM, and CatBoost are all tree-based ensemble models, but they have different design philosophies.</p><ul><li>Random Forest focuses on stability.</li><li>XGBoost focuses on predictive accuracy.</li><li>LightGBM focuses on speed and scalability.</li><li>CatBoost focuses on categorical variables.</li></ul><p>CatBoost’s biggest advantage is that it can use information from categorical variables while reducing the leakage risk that often comes with naive target encoding.</p><p>This makes it especially useful for business data, where categorical variables are often central to the problem.</p><p>If you have mainly used XGBoost or LightGBM so far, CatBoost is worth trying, especially when your data contains many categorical variables.</p><p>You may find that it performs surprisingly well with less preprocessing than you expected.</p><h3>Try CatBoost in Exploratory</h3><p>Exploratory supports CatBoost starting with version 15.5.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*E_nUJ7ou6cg0tRaq.png" /></figure><p>You can build CatBoost models directly from the Analytics view using the UI.</p><p>The basic steps are:</p><ol><li>Open the Analytics view.</li><li>Select CatBoost.</li><li>Choose the target variable.</li><li>Choose the explanatory variables.</li><li>Click Run.</li></ol><p>You can also tune model parameters from the Settings dialog to improve predictive accuracy while controlling overfitting.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*0krs9y2PyFtGbHU2.png" /></figure><p>If you work with data that contains many categorical variables, CatBoost is definitely worth trying.</p><p>You can download the latest version of Exploratory here:</p><p>👉 Download Exploratory</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you do not have an account yet, you can sign up and start a 30-day free trial here:</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>If your trial has already expired, you can launch the latest version and click “Extend Trial” to try it again.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=96c41dc2cd9e" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/introduction-to-catboost-why-it-works-so-well-with-categorical-data-96c41dc2cd9e">Introduction to CatBoost: Why It Works So Well with Categorical Data</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Exploratory v15.5]]></title>
            <link>https://medium.com/learn-dplyr/exploratory-v15-5-c4bdb2540a02?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/c4bdb2540a02</guid>
            <category><![CDATA[statistics]]></category>
            <category><![CDATA[machine-learning]]></category>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[announcements]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Mon, 08 Jun 2026 02:59:37 GMT</pubDate>
            <atom:updated>2026-06-08T02:59:37.568Z</atom:updated>
            <content:encoded><![CDATA[<h3>Exploratory v15.5 Release: Major Analytics Enhancements with CatBoost, Silhouette Method, Proportion Tests, and More</h3><p>We’re excited to announce the release of Exploratory v15.5! 🎉</p><p>Although v15.5 is a minor release in version number, it includes several major enhancements to Exploratory’s Analytics features. This release significantly expands what you can do with machine learning, clustering, and statistical testing.</p><p>In particular, this release focuses on helping users:</p><ul><li>Build more accurate machine learning models.</li><li>Tune machine learning parameters more easily.</li><li>Evaluate clustering results more objectively.</li><li>Run statistical tests that are commonly used in A/B testing, survey analysis, and business analytics.</li></ul><p>As AI increasingly supports data analysis and report generation, it becomes even more important for analysts to understand the models, assumptions, and statistical results behind the analysis.</p><p>Exploratory v15.5 is an important step toward making analytics more powerful, transparent, and trustworthy.</p><p>Here are the highlights.</p><h3>CatBoost Model Added to Machine Learning</h3><p>Exploratory v15.5 adds CatBoost as a new machine learning model.</p><p>CatBoost is a boosting-based machine learning model, similar to XGBoost and LightGBM. In recent years, it has become widely used in the data science community, especially because of its strong handling of categorical variables.</p><p>Business data often contains many categorical variables, such as region, product category, customer type, marketing channel, job role, or industry. CatBoost can be a strong option when your data includes these types of variables.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*LFPsctWhA95nwd9m.png" /></figure><p>Exploratory already supports XGBoost and LightGBM. With the addition of CatBoost, you can now compare multiple high-performance machine learning models and choose the one that works best for your data and use case.</p><p>If you want to improve prediction accuracy, especially with data that includes many categorical variables, we encourage you to try CatBoost.</p><h3>Redesigned Parameter Tuning UI for Machine Learning Models</h3><p>Tuning parameters is an important part of improving machine learning models.</p><p>However, models such as XGBoost, LightGBM, CatBoost, and Random Forest have many parameters. It is not always easy to understand what each parameter does, where to find it, or what value to use.</p><p>In v15.5, we redesigned the parameter setting UI for machine learning models.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*bxUMc1QFUU0HTI-O.png" /></figure><p>Parameters are now grouped by their roles and purposes, and the layout is made more consistent across machine learning models.</p><p>This makes it easier to find parameters related to areas such as:</p><ul><li>Model complexity</li><li>Learning strategy</li><li>Row and column sampling</li><li>Overfitting control</li><li>Validation data and early stopping</li></ul><p>We also added recommended values and practical hints for each parameter.</p><p>For example, you can now see what is likely to happen when a value is set higher or lower. This helps you adjust parameters even if you are not familiar with every technical term.</p><p>Machine learning is not just about building a model once. It is about reviewing the result, adjusting the settings, and improving the model. The redesigned UI makes this tuning process more intuitive and accessible.</p><h3>Silhouette Method Added to K-Means Clustering</h3><p>When using K-Means clustering, you need to decide the number of clusters in advance.</p><p>But choosing the right number of clusters is often difficult.</p><p>Until now, Exploratory supported the Elbow Method to help choose the number of clusters. The Elbow Method is widely used, but sometimes the “elbow” in the chart is not clear, which can make the decision subjective.</p><p>In v15.5, we added the Silhouette Method.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*BCVXykEyUwRGre9l.png" /></figure><p>The Silhouette Method evaluates how well each data point fits within its assigned cluster, while also considering how well it is separated from other clusters.</p><p>The resulting silhouette score helps you evaluate the number of clusters more quantitatively.</p><p>In Exploratory, you can also check not only the average silhouette score, but also the percentage of observations with negative silhouette scores.</p><p>This is important because a high average score alone can be misleading. If many observations have negative silhouette scores, it may indicate that those observations are not assigned to appropriate clusters.</p><p>By looking at both the average silhouette score and the percentage of negative scores, you can make a more realistic decision about the number of clusters.</p><p>We also added silhouette scores by cluster to the summary table.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*7-HQqS9agMEj4WqJ.png" /></figure><p>This allows you to see which clusters are well formed and which clusters may be weaker or less clearly separated.</p><p>Starting with this release, the Silhouette Method is the default method for K-Means Clustering. The Elbow Method is still supported and can be selected from the settings if you want to use it.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*VrEueNfiT9lWzlqX.png" /></figure><h3>Two-Sample Proportion Test Added</h3><p>Exploratory v15.5 adds the Two-Sample Proportion Test.</p><p>This test is used to check whether the proportion of True values is statistically different between two groups.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*huMiIQfnBQUd-oUn.png" /></figure><p>For example, you can use it to answer questions such as:</p><ul><li>Is the conversion rate different between A and B in an A/B test?</li><li>Is the purchase rate different between two customer groups?</li><li>Is the click-through rate different between two campaigns?</li><li>Is the response rate different between two demographic groups?</li></ul><p>Until now, you could use the Chi-Square Test to examine differences in proportions across multiple groups.</p><p>However, when your main goal is to directly compare the proportions between two groups, such as in an A/B test, the Two-Sample Proportion Test is often a more natural fit.</p><p>With this new test, it is now easier to analyze proportion differences in marketing analytics, product analytics, survey analysis, and other common business use cases.</p><h3>One-Sample Proportion Test and One-Sample t-Test Added</h3><p>This release also adds two one-sample tests:</p><ul><li>One-Sample Proportion Test</li><li>One-Sample t-Test</li></ul><p>These tests are used to check whether the proportion or average in your sample is statistically different from a predefined benchmark, target, or hypothesized value.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*iMu7gJd3ewefqIu9.png" /></figure><p>For example, with the One-Sample Proportion Test, you can answer questions such as:</p><ul><li>Is the conversion rate higher than the target value of 10%?</li><li>Is the satisfaction rate different from the expected value of 80%?</li><li>Is the defect rate higher than the acceptable threshold of 5%?</li></ul><p>With the One-Sample t-Test, you can answer questions such as:</p><ul><li>Is the average revenue higher than the target value?</li><li>Is the average satisfaction score different from the benchmark?</li><li>Is the average processing time shorter than the expected value?</li></ul><p>These tests make it possible to evaluate your data against a target or benchmark, not only against another group.</p><h3>Note: Table Column Width Can Now Be Adjusted</h3><p>We also made an improvement to Notes.</p><p>In the Note Editor, you can use tables to organize your analysis results and explanations. With this release, you can now adjust table column widths with your mouse, and the adjusted widths are preserved when the note is published.</p><p>Note Editor:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*VvJPSNbDP4o77Cnn.png" /></figure><p>Published Note:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*fbSylNpXepE50XMl.png" /></figure><p>This makes it easier to improve the readability of tables and create cleaner, more polished analytical reports.</p><h3>Summary of New Features in Exploratory v15.5</h3><p>Here are the main additions and improvements in Exploratory v15.5.</p><ul><li>Added CatBoost as a new machine learning model.</li><li>Redesigned the parameter setting UI for machine learning models.</li><li>Added the Silhouette Method to K-Means clustering.</li><li>Added silhouette scores by cluster to the clustering summary.</li><li>Added the Two-Sample Proportion Test.</li><li>Added the One-Sample Proportion Test.</li><li>Added the One-Sample t-Test.</li><li>Improved Notes so table column widths can be adjusted and preserved after publishing.</li></ul><p>This release also includes many bug fixes and other improvements. For details, please see the <a href="https://exploratory.io/release-notes">release notes</a>.</p><h3>Try Exploratory v15 !</h3><p>Exploratory v15 includes many new features designed to support trustworthy data analysis in the age of AI.</p><p>In v15.0, we introduced major improvements to make data wrangling workflows more visible, explainable, documented, and easier to fix when errors occur.</p><p><a href="https://blog.exploratory.io/exploratory-v15-the-3-rs-for-trustworthy-data-analysis-in-the-ai-era-172dab64d07b">Exploratory v15 - The 3 Rs for Trustworthy Data Analysis in the AI Era</a></p><p>With v15.5, we are making major improvements to Analytics.</p><p>You can now not only prepare clean and reproducible data, but also use machine learning, clustering, and statistical testing to gain deeper insights from that data.</p><p>As AI increasingly helps with analysis and report writing, it becomes more important for analysts to understand how models are built, how parameters affect results, and how statistical results should be interpreted.</p><p>We believe Exploratory v15.5 is an important milestone toward building a trustworthy analytics foundation for the AI era.</p><p>Please download the latest version and try it out.</p><p>👉 Download Exploratory</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an account yet, you can start a 30-day free trial here.</p><p>👉 Start Free Trial</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>That’s all for today.</p><p>If you have any feedback about Exploratory v15, please feel free to comment below!</p><p>We will continue working hard to improve Exploratory and make data analysis more trustworthy and accessible in the AI era.</p><p>Thank you!</p><p>Kan,<br>CEO/Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=c4bdb2540a02" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/exploratory-v15-5-c4bdb2540a02">Exploratory v15.5</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Exploratory v15 — The 3 Rs for Trustworthy Data Analysis in the AI Era]]></title>
            <link>https://medium.com/learn-dplyr/exploratory-v15-the-3-rs-for-trustworthy-data-analysis-in-the-ai-era-172dab64d07b?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/172dab64d07b</guid>
            <category><![CDATA[announcements]]></category>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[ai]]></category>
            <category><![CDATA[data-wrangling]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Mon, 08 Jun 2026 02:52:31 GMT</pubDate>
            <atom:updated>2026-06-08T02:52:31.569Z</atom:updated>
            <content:encoded><![CDATA[<h4>Readability, Reproducibility, and Reliability</h4><p>I hope you are doing well.</p><p>This is Kan from Exploratory.</p><p>We are very excited to announce the release of Exploratory v15! 🎉</p><p>The theme of this release is:</p><p><strong>“The 3 Rs for Trustworthy Data Analysis in the AI Era”</strong></p><ul><li><strong>Readability</strong> — Data transformation processes should be easy to read and understand.</li><li><strong>Reproducibility</strong> — Analysis results should be reproducible.</li><li><strong>Reliability</strong> — Reproducibility can be maintained even if the data &amp; structure change.</li></ul><p>Today, AI is making it easier than ever to generate charts, dashboards, reports, and even analytical insights automatically.</p><p>However, no matter how beautifully a dashboard or analysis result is generated, if we don’t understand how the underlying data was processed and transformed, it becomes difficult to trust the result with confidence.</p><p>Questions like:</p><ul><li>“Where did this number come from?”</li><li>“How transformations, aggregations, and filters were applied?”</li><li>“Can this analysis still be reproduced with the new data?”</li></ul><p>are becoming more important than ever in the AI era.</p><p>To address this challenge, Exploratory v15 introduces new capabilities designed to <strong>improve the readability, reproducibility, and reliability of data preparation and analysis workflows.</strong></p><p>Here are some of the major new features in v15:</p><ul><li><strong>Step Diagram</strong> to visualize data transformation pipelines</li><li><strong>AI Step Summary</strong> to explain data transformation steps in natural language</li><li><strong>Automatic Document Generation</strong> for the entire data preparation process</li><li><strong>AI-Powered Auto Fix</strong> for data transformation errors</li><li><strong>AI Summaries</strong> for Charts and Analytics</li><li>Enhancements to Summary View, Charts, Tables, and Performance</li></ul><p>Below is a quick overview of the new features in Exploratory v15.</p><h3>Step Diagram: Visualizing Data Transformation Pipelines</h3><p>Exploratory v15 introduces the <strong>Step Diagram</strong>, which allows you to visually trace the flow of data transformations from charts, dashboards, and reports all the way down to source data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*g-FKpYA68d2ShxEn" /></figure><p>You can now easily understand:</p><ul><li>Where the data is sourced from</li><li>How the data is transformed and lead to specific charts or KPIs</li><li>Where and how data are combined (join &amp; merge)</li><li>Where data frame was branched off</li></ul><p>The Step Diagram can also be accessed directly from charts and numbers (metrics) inside dashboards and notes.</p><p><strong>Dashboard</strong></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*f1o_3uXoIz901i-_" /></figure><p><strong>Note</strong></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*ZZP_O_cT4s_cM8eV" /></figure><p>This makes it much easier to answer questions like:</p><p>“How exactly was this number created?”</p><h3>AI Step Summary</h3><p>We also added <strong>AI-generated Step Summaries</strong>, which explain data transformation steps in natural language.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*lZCpJR8qmhod7HBa" /></figure><p>For example:</p><ul><li>“Filling NA values with previous values after sorting data by purchase date”</li><li>“Joining customer data with order data using Customer ID”</li><li>“Aggregating sales by region”</li><li>“Keeping only customers with more than two purchases”</li></ul><p>This makes data transformation workflows easier to understand, even for users who are not familiar with underlying code or detailed transformation logic.</p><h3>Automatic Documentation Generation for Data Preparation</h3><p>Exploratory v15 can now automatically generate documentation for the entire data preparation process.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*SIN4_XfOcg-z9Odx" /></figure><p>Using AI-generated summaries, Exploratory creates a structured document that includes:</p><ul><li>An overview of the entire workflow</li><li>Descriptions of each transformation step</li><li>Organize by grouped sections based on transformation purpose</li></ul><p>The generated document can then be saved directly as an Exploratory Note, allowing you to edit, expand, and share it with your team.</p><p>This is especially useful for:</p><ul><li>project handoffs</li><li>team reviews</li><li>supplemental documentation for reports</li><li>explaining analysis logic</li><li>auditing and validation of data preparation workflows</li></ul><p>Until now, dashboards and notes could be shared easily, but the context behind the data preparation process was often difficult to communicate because documenting it manually takes significant time.</p><p>With automatic documentation generation, you can now share not only the final result, but also the context behind how the result was created.</p><h3>AI-Powered Auto Fix for Data Transformation Errors</h3><p>From the very beginning of Exploratory, reproducibility has been one of the core design principles.</p><p>Exploratory records data transformations as step-by-step workflows (at the right-hand side) and automatically manages relationships between:</p><ul><li>transformation steps</li><li>data frames</li><li>charts</li><li>analytics</li><li>dashboards and notes</li></ul><p>This enables reproducible data analysis workflows.</p><p>However, maintaining reproducibility is not always easy because real-world data changes constantly.</p><p>For example:</p><ul><li>column names changed in the source file</li><li>values are presented in a different format, which causes data types change</li><li>renamed columns or data type changes break downstream transformations</li></ul><p>To help address these problems, Exploratory v15 introduces <strong>AI-powered Auto Fix for data transformation errors</strong>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*jun_KR3g44jjwHG1" /></figure><p>Rather than simply analyzing the error message itself, the AI reviews the surrounding data transformation context and suggests practical fixes while preserving the integrity of the transformation pipeline.</p><p>For example, it can:</p><ul><li>find similar column names and update column name references</li><li>add appropriate data type conversion steps</li><li>disable the step that causes the downstream error in the pipeline</li><li>fix syntax errors in calculations</li></ul><p>Users can review the suggested fixes and apply the appropriate fix.</p><p>This makes it easier to maintain reproducible and reliable data transformation pipeline even when there are accidental changes in source data or in the data transformation steps.</p><h3>AI Summaries for Charts and Analytics</h3><h4>Analytics</h4><p>Analytics view in Exploratory automatically generates analysis reports, but analytical outputs often contain a large amount of information (or, too much info!).</p><p>AI Summary helps users quickly identify the most important findings before diving into the detailed report.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*NRcmaJzk6XD7HudR" /></figure><p>This makes interpreting analysis results much more efficient.</p><h4>Charts</h4><p>AI Summary for Charts explains patterns, trends, and statistical signals in an easy-to-understand way.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*tVBfQHlXmVfOX8kD" /></figure><p>For example, with XmR charts, AI Summary can automatically identify whether special-cause variation signals are present beyond expected variation.</p><p>This helps users notice important insights that might otherwise be overlooked.</p><h3>Enhancements to Summary View</h3><p><strong>Interactive Filtering</strong></p><p>You can now click bars directly in Summary View charts to keep or exclude related records interactively.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*bga9unz6mCCDulw4" /></figure><h3>Outlier Analysis</h3><p>We also added a new <strong>Outlier Analysis</strong> feature to Summary View.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*n18yBRW7W7GoHT4W" /></figure><p>This allows users to quickly visualize how outliers in one variable relate to other variables in the dataset.</p><h3>Custom Charts</h3><p>Exploratory v15 introduces <strong>Custom Charts</strong>, allowing users to write R code directly inside Chart View to create highly customizable visualizations.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*BFr_6CJ81n_zWMiM" /></figure><p>By leveraging R’s rich visualization packages, users can now create charts beyond standard built-in chart types directly within Exploratory.</p><h3>Drag-and-Drop Column Width Adjustment</h3><p>You can now resize columns in pivot tables, summary tables, and regular tables using drag-and-drop.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*Vh-cPW4s0mL5OrFP" /></figure><p>This makes it easier to improve readability and optimize table layouts for dashboards and reports.</p><h3>Performance Improvements</h3><p>Exploratory v15 also includes major performance improvements.</p><p>Key areas include:</p><ul><li>Faster Exploratory startup</li><li>Faster loading for Notes and Dashboards</li></ul><p>Startup performance improvements are especially noticeable for users with many projects.</p><p>Additionally, Exploratory now upgrades its R foundation to <strong>R 4.5</strong> and R packages.</p><p>One notable package upgrade is the <strong>dplyr v1.2</strong>, which significantly improves the performance of functions such as:</p><ul><li>case_when()</li><li>if_else()</li></ul><p>These functions are heavily used in Exploratory’s conditional calculation steps.</p><h3>Try Exploratory v15</h3><p>In the AI era, trustworthy analysis requires understanding how data was created, transformed, and analyzed.</p><p>Exploratory v15 is designed to support exactly that.</p><p>We hope you will try the latest version and experience the new capabilities yourself.</p><p>👉 <strong>Download Exploratory</strong><br>​<a href="https://exploratory.io/download">Download Exploratory</a></p><p>If you do not yet have an account, you can start a 30-day free trial.</p><p>👉 <strong>Start Your Free Trial</strong><br>​<a href="https://exploratory.io/">Start Free Trial</a></p><p>That’s all for today.</p><p>If you have any feedback about Exploratory v15, please feel free to comment below!</p><p>We will continue working hard to improve Exploratory and make data analysis more trustworthy and accessible in the AI era.</p><p>Thank you!</p><p>Kan,<br>CEO/Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=172dab64d07b" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/exploratory-v15-the-3-rs-for-trustworthy-data-analysis-in-the-ai-era-172dab64d07b">Exploratory v15 — The 3 Rs for Trustworthy Data Analysis in the AI Era</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[In the Age of AI, R’s Readability Becomes a Superpower]]></title>
            <link>https://medium.com/learn-dplyr/in-the-age-of-ai-rs-readability-becomes-a-superpower-4e9b59beeabd?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/4e9b59beeabd</guid>
            <category><![CDATA[rstats]]></category>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[ai]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Thu, 23 Apr 2026 16:09:29 GMT</pubDate>
            <atom:updated>2026-04-23T18:06:40.302Z</atom:updated>
            <content:encoded><![CDATA[<h4>How readable code enables reliable, reproducible, and interactive data analysis when AI writes the code</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*2leH0LOrAPM4cq9E.png" /></figure><p><em>Originally published on Exploratory. — </em><a href="https://exploratory.io/note/kanaugust/jmq2Oyh4Jv"><em>Link</em></a>.</p><p>We are entering a new phase of data science.</p><p>You don’t write most of the code anymore, but AI does.</p><p>At first glance, this feels like a liberation. For years, learning data science meant learning how to code in Python, R, SQL, etc. Now, you can simply describe what you want, and the code gets generated for you.</p><p>This makes you wonder…</p><blockquote><em>If AI can generate the code, do we still need to learn programming like R or Python at all for data science?</em></blockquote><p>Well, it depends…</p><h3>The Real Problem: Answers without Trust</h3><p>In many areas of machine learning we can evaluate our results against something concrete. There is a ground truth. A label.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*lHvocHFXcIc1M2j6.png" /></figure><p>And we can evaluate how well we are doing with a set of well-defined metrics.</p><p>For this type of data science, you may never have to look at the underlying code that builds prediction models, so what’s the point of learning how to code? You may ask.</p><p>However, when it comes to <strong>data analysis side of data science</strong> things are very different.</p><p>Because, there is no clear pre-defined answer that we can evaluate against at the end.</p><p>You are not verifying result against a particular answer, instead you are building an answer. You are exploring data in order to understand something that was previously unknown.</p><p>And this makes a difference between generating code for <strong>data analysis</strong> side of data science and for <strong>building prediction models</strong> side of data science.</p><p>In fact, the real difficulty in data analysis has always been how you build the answers you can trust, not how to write code.</p><p>In order to trust the result, you have to have trust in the process that produced the result. You have to be confident that each data transformation such as filter, calculations, aggregation, combine with other data, cleaning text data, etc. is doing exactly what you intended.</p><p>Not approximately.</p><p>“Close enough” is not enough in data analysis or building metrics or KPIs, it has to be ‘Exact and Accurate’ because people (or even agents!) will make decisions based on the result.</p><h3>Readability, the Most Important Part of Programming Language</h3><p>This leads to an uncomfortable but unavoidable truth:</p><blockquote><em>If you cannot read the code, you cannot trust the result.</em></blockquote><p>This is where <strong>Readability</strong> becomes the most critical feature when you choose programming language for data analysis.</p><p>In the context of AI generated code, people often talk about the accuracy of the code, how fast the performance, what is the cost, which AI model to use, etc.</p><p>But in the context of data analysis, none of these are as important as <strong>Readability.</strong></p><h3>From Readability to Trust</h3><p>Before making further discussion, I want to make it clear.</p><p>When I say Readability, it is not just about aesthetics or personal preference. It is not about whether code looks “clean” or “elegant.”</p><p>It is about whether you can understand what the code is actually doing.</p><p>When AI generates code, it is easy to <strong>assume</strong> that the result is correct. But small details matter. A slightly different grouping, an unintended filter, a silent drop of missing values, any of these can change the outcome in meaningful ways.</p><p>If the code is readable, you can pause and inspect it. You can follow the transformation step by step and confirm that it aligns with your intention. The code becomes something you can reason about.</p><p>Here’s an example of R code that provides high readability.</p><pre>sales %&gt;%  <br>   filter(year &gt;= 2020) %&gt;%  <br>   group_by(region) %&gt;%  <br>   summarize(total_sales = sum(sales))</pre><p>It can be read as:</p><p>“Filters data for 2020 or later, and calculate total sales for each group of region.”</p><p>It’s readable, and you can understand what it’s trying to do.</p><p>If it is not readable, you are left with a different experience. You are not verifying the result, you are just trusting it faithfully. And in data analysis, blind trust is not a viable strategy.</p><p>Trust, then, is not something AI gives you.</p><p>It is something you build through <strong>readability</strong>.</p><h3>From Readability to Reproducibility</h3><p>But this is only the beginning.</p><p>In a business context, data analysis rarely happens only once. Data changes. Questions evolve. The same analysis needs to be run again next week, next month, or with slightly different conditions.</p><p>At that point, the question becomes:</p><p>“Can you reuse what you already did before?”</p><p>If you understand the code, the answer is simple. You run it again. And if needed, you adjust a condition, extend a calculation, etc. The original logic in data transformation and analysis remains intact, and the analysis evolves naturally over time.</p><p>If you don’t, it breaks.</p><p>You go back to AI. You describe the task again. The AI generates something similar, but may not be identical. Small differences creep in. Over time, you accumulate multiple versions of “the same” analysis that are, in fact, not quite the same.</p><p>This is the problem when you don’t have reproducible data analysis pipeline. And if you can’t reproduce you’ll lose trust quickly.</p><p>Again, the question is not whether you or AI can write, but it is whether the code, either generated by AI or written by you, is readable enough to be reused with confidence.</p><h3>From Readability to Interactivity</h3><p>There is one more dimension, and it is perhaps the most subtle.</p><p>Data analysis is not a linear process. It is not a sequence of predefined steps. It is a iterative process, a back-and-forth between you and the data, as if you are having a conversation with data.</p><p>You might begin with something simple:</p><p>“Show me sales by year.”</p><p>But almost immediately, your thinking moves forward.</p><p>-&gt; What about monthly trends?</p><p>-&gt; What about regional differences?</p><p>-&gt; What about repeat customers?</p><p>-&gt; Is there a correlation with marketing spend?</p><p>Each question generates an answer, which generates a new question, which generates an answer and so on. It’s an iterative process and the center of the process is you as a curious and creative being.</p><p>And you want this process to be fast, as fast as the speed of thought.</p><p>But when every step requires writing a prompt, waiting for AI to respond, reading the generated code, and then inspecting whether the result is correct, that flow is interrupted. The conversation slows down.</p><p>You are no longer exploring, you are coordinating with a tool.</p><p>However, something different happens when the code is readable.</p><p>You stop relying on prompts for every step. Instead, you begin to adjust the code directly. A grouping changes. A filter is added. A calculation is modified. The feedback loop tightens.</p><p>Now, the pace of analysis starts to match the pace of your thinking.</p><p>That is what true exploratory data analysis feels like.</p><p>It’s interactive and iterative.</p><h3>The Shift We Often Miss</h3><p>It is tempting to think that AI reduces the importance of programming.</p><p>In reality, it does the opposite.</p><p>It shifts the role of programming from <strong>writing code</strong> to <strong>understanding code</strong>, at least in a context of data analysis.</p><p>And once that shift happens, the most important property of programming language is no longer what it can do.</p><p>It is how easily you can read what it does.</p><p>Because that is what determines whether you can:</p><ul><li><strong>trust</strong> the result</li><li><strong>reproduce</strong> the process</li><li>and <strong>interact</strong> with the data in a meaningful way</li></ul><p>And all of that begins in the same place.</p><p><strong>Readability.</strong></p><h3>3 Technical Foundations of R That Enables Readability</h3><p>If readability is the key to trust, reproducibility, and interactivity, then the next question is:</p><p>Which data science language provides high readability?</p><p>And the short answer answer is R.</p><p>R provides by far the most readable code. When I say R, I don’t necessary mean the base R, but R with ‘tidyverse’, which is a set of packages that most of the modern R practitioners use.</p><p>R’s tidyverse code resembles the description of the data operation in English as we have seen above.</p><pre>sales %&gt;%  <br>   filter(year &gt;= 2020) %&gt;%  <br>   group_by(region) %&gt;%  <br>   summarise(total_sales = sum(sales))</pre><p>But, the tidyverse didn’t become highly readable by itself. It was R’s unique foundation that enables it.</p><p>Clause Wilke has articulated this so well in his blog post when he talks about why R is better than Python.</p><ul><li><a href="https://blog.genesmindsmachines.com/p/python-is-not-a-great-language-for-2e0">Python is not a great language for data science. Part 2: Language features</a></li></ul><p>And among many reasons listed in the post, I think the following 3 characteristics of R are especially important for enabling R to be highly readable.</p><ol><li>Separation of <strong>Logic vs Logistics</strong></li><li>Built-in <strong>Vectorization</strong></li><li>Support for <strong>Non-Standard Evaluation (NSE)</strong></li></ol><p>These three work together to determine whether code feels like a clear expression of thought or a set of instructions for a machine.</p><p>In order to highlight these unique features, as Clause does in his blog post, I’m going to compare R against Python. Python is often used by most of the AI models when you ask data science / analysis related questions without specifying your preference of language.</p><h3>1. Separation of Logic vs Logistics</h3><p>Every piece of data analysis code contains two layers:</p><ul><li><strong>Logic</strong> → What you want to do</li><li><strong>Logistics</strong> → How the computer executes it</li></ul><p>For example:</p><blockquote><em>“Calculate total sales by region for 2020 and after”</em></blockquote><p>This is <em>logic</em>.</p><p>But implementing it may involve:</p><ul><li>loops</li><li>indexing</li><li>temporary variables</li><li>restructuring outputs</li></ul><p>This is <em>logistics</em>.</p><p><strong>Example with R</strong> <strong>(Tidyverse)</strong></p><p>Now, here is how you write in order to achieve the above with R (tidyverse).</p><pre>sales %&gt;%  <br>   filter(year &gt;= 2020) %&gt;%  <br>   group_by(region) %&gt;%  <br>   summarise(total_sales = sum(sales))</pre><p>This is almost a direct translation of the question:</p><ul><li>filter year &gt;= 2020</li><li>group by region</li><li>summarize total sales</li></ul><p>You see only the logic, almost <strong>no visible logistics</strong></p><h3>Example with Python (Panda)</h3><p>Now, here is an example code with Python with Panda library to do the same.</p><pre>sales[sales[&#39;year&#39;] &gt;= 2020] \<br>    .groupby(&#39;region&#39;) \<br>    .agg(total_sales=(&#39;sales&#39;, &#39;sum&#39;)) \<br>    .reset_index()</pre><p>Still readable, but notice the logistics creeping in:</p><ul><li>sales[&#39;year&#39;] → how to access column</li><li>agg(...) → syntax detail</li><li>(&#39;sales&#39;, &#39;sum&#39;) → implementation detail</li><li>.reset_index() → structural fix</li></ul><p>These are not part of the <strong>question</strong>. They are part of the <strong>machinery</strong>.</p><p>If AI decided to write Python code without the panda library, this is how it would look like. You can see the true nature of Python:</p><pre>result = {}<br><br>for row in sales:    <br>    if row[&#39;year&#39;] &gt;= 2020:        <br>       region = row[&#39;region&#39;]   <br>       if region not in result:     <br>           result[region] = 0  <br>       result[region] += row[&#39;sales&#39;]</pre><p>Now we are fully in <strong>logistics mode</strong>:</p><ul><li>Manual loop</li><li>Dictionary initialization</li><li>Key checks</li><li>Accumulation logic</li></ul><p>This code does the same as the above R example, but it forces you to think like a machine.</p><p><strong>Why This Difference Exists</strong></p><p>This difference is not accidental. It comes from the design philosophy of the languages.</p><p>R (tidyverse) is designed for expressing <strong>intent</strong>. It uses <strong>data-centric grammar (verbs like filter, mutate, summarize, etc.)</strong> And this is only possible because the following features of R, which I’m going to talk about in the next sections.</p><ul><li>Built around <strong>vectorized operations</strong></li><li>Supports <strong>non-standard evaluation (NSE)</strong></li></ul><p>On the other hand, Python is designed for general-purpose programming. It’s built around <strong>objects and structures</strong>, and it requires explicit references like (df[&#39;column&#39;]).</p><p>It lacks native data-manipulation grammar. And even with a library like pandas, Python still carries its general-purpose DNA.</p><p>When logistics dominates,</p><ul><li>Code becomes harder to read</li><li>Verification becomes harder</li><li>Modification becomes slower</li></ul><p>In R code, logic dominates. And this makes R code closer to human thinking.</p><h3>2. Built-in Vectorization</h3><p>Now let’s go one level deeper.</p><p>Even if you separate logic from logistics, you still need a way to apply that logic to data.</p><p>This is where <strong>vectorization</strong> comes in.</p><p>With vectorization, you can apply an operation to an entire column at once, instead of looping row by row.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*2EhQkxgElpdVKy4e.png" /></figure><p>Let’s say you want to make a simple calculation like “multiple sales by 1.1”.</p><p>With R, you can write something like this.</p><pre>sales * 1.1</pre><p>This applies to all values in the sales column (or variable).</p><p>No loop. No special function. Just the logic.</p><p>But in Python, you must iterate explicitly.</p><pre>[s * 1.1 for s in sales]</pre><p>With the pandas or the NumPy library it could be close to R.</p><pre>df[&#39;sales&#39;] * 1.1</pre><p>In Python, vectorization is not part of the language itself unlike R. Instead, it is outsourced to libraries. Therefore, you need to deal with the difference among the libraries:</p><ul><li>Python lists → not vectorized</li><li>NumPy arrays → vectorized</li><li>pandas Series → vectorized (differently)</li><li>Polars Series → vectorized (again differently)</li></ul><p>Each has:</p><ul><li>different APIs</li><li>different behaviors</li><li>different edge cases</li></ul><p>So you end up writing with the numpy package:</p><pre>import numpy as np<br><br>sales = np.array(sales)<br>sales * 1.1</pre><p>Or, with the panda package:</p><pre>pd.Series(sales) * 1.1</pre><p>And this gets worse when you want to calculate conditionally.</p><p>Let’s say if you want to calculate something like this.</p><blockquote><em>“If sales &gt; 100, apply 10% discount”</em></blockquote><p>In the case of R, it’s pretty simple.</p><pre>ifelse(sales &gt; 100, sales * 0.9, sales)</pre><p>You have the following three.</p><ul><li>condition</li><li>calculation</li><li>fallback</li></ul><p>But in the case of Python, it becomes convoluted.</p><pre>[s * 0.9 if s &gt; 100 else s for s in sales]</pre><p>Now:</p><ul><li>explicit iteration</li><li>inline conditional logic</li><li>more cognitive load</li></ul><p>Even with pandas package,</p><pre>df[&#39;sales&#39;] = np.where(df[&#39;sales&#39;] &gt; 100, df[&#39;sales&#39;] * 0.9, df[&#39;sales&#39;])</pre><p>It’s better, but still:</p><ul><li>referencing columns repeatedly</li><li>using external function (np.where)</li><li>mixing libraries (pandas + NumPy)</li></ul><h3>Why this matters</h3><p>Vectorization is not just about performance. It’s about <strong>thinking at the right level of abstraction</strong>.</p><p>When vectorization is built into the language:</p><ul><li>You think in <strong>columns and transformations</strong></li></ul><p>When it is not:</p><ul><li>You fall back to <strong>rows and loops</strong></li></ul><h3>3. Non-Standard Evaluation (NSE)</h3><p>Now we reach the most important layer.</p><p>Even if you can think in logic and can operate on vectors, you still need a way to <strong>express your idea naturally</strong>.</p><p>This is where NSE (Non-Standard Evaluation) comes in.</p><h3>What is Non-Standard Evaluation?</h3><p>NSE (Non-Standard Evaluation) means that you can write code that <em>looks like you are directly operating on the data</em>, even though the data lives inside a data frame.</p><p>In simpler terms, you can refer to columns as if they were normal variables.</p><p>R has NSE support, so you can specify the column (variable) names in functions like this.</p><pre>df %&gt;%<br>  mutate(profit = sales - cost)</pre><p>You can read this code literary as:</p><ul><li>profit equals sales minus cost</li></ul><p>No quotes, no indexing, and no function around the column name. You simply type column names as if you write in a report.</p><p>On the other hand, Python doesn’t have NSE, hence the column (variable) needs to be specified as below.</p><pre>df[&#39;profit&#39;] = df[&#39;sales&#39;] - df[&#39;cost&#39;]</pre><p>Notice that</p><ul><li>Columns must be accessed as strings (&#39;sales&#39;)</li><li>Repeated df[...]</li></ul><p>This may seem small, but it becomes worse as we expand.</p><p><strong>Example 1: Grouping and Summarizing</strong></p><p>Let’s say you want to:</p><blockquote><em>“Group by region, calculate average sales”</em></blockquote><p>In R, you will write like the below.</p><pre>df %&gt;%<br>  group_by(region) %&gt;%<br>  summarise(avg_sales = mean(sales))</pre><p>This can be easily read as ‘group by region, calculate average sales’.</p><p>But with Python, you will need to write like this.</p><pre>df.groupby(&#39;region&#39;)[&#39;sales&#39;].mean().reset_index(name=&#39;avg_sales&#39;)</pre><ul><li>&#39;region&#39;, &#39;sales&#39; → string references</li><li>chaining syntax tied to structure</li><li>.reset_index() → logistics</li></ul><p>This is when non-programmers start losing it.</p><p>And it gets even worse.</p><p><strong>Example 2: Sorting based on Calculation</strong></p><p>Let’s push further. If you want to do something like:</p><blockquote><em>“Sort by cosine of sales”</em></blockquote><p>With <strong>R</strong>, it’s pretty simple.</p><pre>df %&gt;%<br>  arrange(cos(sales))</pre><p>With <strong>Python</strong>, not so much.</p><pre>import numpy as np<br><br>df.assign(<br>    cos_sales=lambda d: np.cos(d[&#39;sales&#39;])<br>).sort_values(&#39;cos_sales&#39;) \ <br> .drop(columns=[&#39;cos_sales&#39;])</pre><p>Now we see:</p><ul><li>temporary column</li><li>lambda function</li><li>repeated references</li><li>cleanup step</li></ul><p>We’re not comparing the capability between R and Python. Both codes can do the same job. But which one would you be more comfortable when you are accountable for the produced result?</p><p>NSE turns code into a language for expressing ideas, rather than constructing the computational instruction.</p><h3>Bringing It All Together</h3><p>These three principles are not independent. They reinforce each other.</p><p>When all three are present,</p><ul><li>You express <strong>logic</strong>, not logistics</li><li>You operate on <strong>data as a whole</strong></li><li>You write <strong>expressions, not instructions</strong></li></ul><p>The result is the code that reads like the question you are asking.</p><h3>The Shift in the AI Era</h3><p>In the past, programming was about writing code. And because humans wrote that code, readability mattered, primarily as a matter of maintainability and collaboration.</p><p>But today, we live in a different time.</p><p>AI writes a significant portion of the code we use, and that changes the role of the human entirely.</p><p>You are no longer the author or even the maintainer of the code. You are the reviewer of it.</p><p>You didn’t write it. You didn’t think through each line as it was constructed. And yet, you are responsible for deciding whether the result produced by the code is correct.</p><p>That responsibility cannot be delegated to AI or others.</p><p>And this is where the hierarchy of importance flips.</p><p>Readability is no longer a “nice-to-have.” It becomes the foundation upon which everything else depends.</p><p>Because only when code is readable can you:</p><ul><li>verify its correctness</li><li>trust the result</li><li>reuse the logic</li><li>and modify it to satisfy ever evolving questions you have.</li></ul><p>Reliability, reproducibility, and interactivity are not separate concerns.</p><p>They are all downstream of one thing.</p><p>That’s <strong>Readability.</strong></p><h3>Pick the Language That Gives You Readability</h3><p>If your goal is to work effectively with AI-generated code, you need a language where:</p><ul><li>the code reflects the logic directly</li><li>transformations are easy to follow</li><li>small changes can be made without friction</li></ul><p>This is where R, especially with the tidyverse, stands out.</p><p>Not because it can do something others cannot.</p><p>But because it expresses ideas in a way that is intuitively understandable.</p><p>When AI generates R code, what you see is not machinery, you see intent.</p><p>And that makes all the difference.</p><h3>Why This Matters for Exploratory</h3><p>By the way, this perspective is not something we arrived at recently.</p><p>It is, in fact, the original reason we built Exploratory in the first place.</p><p>About 10 years ago, when we started Exploratory, we made a deliberate decision: build UI on top of R and the tidyverse.</p><ul><li><a href="https://blog.exploratory.io/introducing-exploratory-desktop-ui-for-r-895d94ef3b7b">Introducing Exploratory Desktop — UI for R</a></li></ul><p>At the time, this was not the most obvious choice to many. But we believed something that, in hindsight, has only become more true over time.</p><p>Data analysis is not just about capability.</p><p>It is about how easily humans can <strong>understand</strong> and <strong>work with</strong> that capability.</p><p>We saw that the R &amp; dplyr (a part of tidyverse) offered something unique.</p><p>Not just a set of powerful functions, but a way of expressing data transformations that was readable, consistent, and aligned with how we think. That readability, combined with flexibility, made it ideal for exploratory data analysis.</p><p>So we built an UI experience on top of it.</p><p>The goal was simple: to make the power of R accessible to non-technical users without a need of writing code.</p><p>That was the beginning of Exploratory.</p><h3>What is the bottleneck for data analysis today?</h3><p>10 years later, things have changed dramatically.</p><p>It is now 2026, and we are deep in the era of AI and LLM (large language models). Writing code is no longer the hard part. AI can write R, Python, SQL — almost instantly and entirely.</p><p>And this begs a question.</p><p>If writing code is no longer the bottleneck, then what is?</p><p>We started asking our customers,</p><blockquote><em>What </em>is<em> the bottleneck for data analysis today in the era of AI?</em></blockquote><p>And the answers were surprisingly consistent.</p><p>The bottlenecks were:</p><p><strong>Reliability</strong> and <strong>Reproducibility</strong>.</p><p>This brings us back to the core idea of this post.</p><p>If reliability and reproducibility are the problems, then how can we solve?</p><p>And the answer, as we’ve seen, is <strong>readability</strong> that is supported by the right interface — AI prompt and UI.</p><p>This is where Exploratory fits in, especially in the AI era.</p><h3>Data Analysis in Exploratory</h3><p>In Exploratory, you can type in how you want to transform and prepare data in the prompt UI, then AI generates R (tidyverse) code. Not as something hidden behind the scenes, but as something you can see, read, and understand.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*q_wAzdGy_ArbX8uV.png" /></figure><p>You can read the code with a help of the code description to see what it does and what you can expect from it.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*wiVxmoq5mSCCFgy-.png" /></figure><p>But it doesn’t disappear into somewhere it’s hard to find like Excel where the calculations and transformations are spread across a bunch of cells.</p><p>It becomes part of a visible, inspectable data wrangling step.</p><p>Each step is preserved as part of a pipeline — a sequence of data transformations that you can revisit, modify (such as changing the data source files or underlying SQL queries, updating calculations, changing the data filter conditions, etc.), and rerun at any time.</p><p>You can see the updated data visualized by Summary view where each column shows its summary information with automatically generated chart and metrics.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*c1GlX6-fm36-bqtT.png" /></figure><p>Or, simply use Chart view where you can visualize the data with various chart types and chart features with point-and-click.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*xu9DkU8iTEANUK9u.png" /></figure><p>This makes a powerful data exploratory experience, which is highly visual, iterative, and interactive.</p><p>Over time, this changes how you work.</p><p>You are no longer repeatedly asking AI to regenerate the same logic. You are refining and evolving an analysis that you understand.</p><p>When new data arrives, you don’t start over. You update the source step and rerun the process.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*JSIVGV9yozEMioWJ.png" /></figure><p>When a new question emerges, you don’t need to start from scratch. You add a new step and continue the flow.</p><p>At this point, your analysis becomes:</p><ul><li><strong>Reliable</strong>, because you can verify every step</li><li><strong>Reproducible</strong>, because the process is preserved</li><li><strong>Interactive</strong>, because you can move at the speed of your thinking</li></ul><p>And this is not simply because AI is generating code.</p><p>It is because that code is <em>readable</em>, and because the environment is designed to support that readability.</p><h3>Final Thought</h3><p>If there is one idea to take away from all of this, it is this:</p><p>The biggest shift in the AI era is not <strong>who writes the code</strong>. It is <strong>who is responsible for understanding the code</strong> AI generates in the world of Data Analysis.</p><p>In the past, you wrote the code, so you naturally understood it.</p><p>Today, AI writes the code.</p><p>Which means understanding no longer comes for free.</p><p>It becomes something you must actively regain.</p><p>And the only way to do that is through <strong>readability</strong>.</p><p>When code is generated automatically, the value no longer lies in producing it.</p><p>The value lies in being able to <strong>read it instantly, verify it confidently, and adapt it flexibly</strong>.</p><p>That is why programming itself hasn’t become less important.</p><p>It has simply changed its role.</p><p>From:</p><blockquote><em>writing code</em></blockquote><p>To:</p><blockquote><em>understanding and validating code</em></blockquote><p>Therefore, in this AI era, the choice of language is not about capability alone, but it’s about readability.</p><p>It is about whether you can understand what has been generated quickly.</p><p>Because that is what helps you trust the result.</p><p>That is what makes it reusable.</p><p>That is what makes it interactive.</p><p>And that is why R, especially tidyverse, feels increasingly aligned with the future of data science.</p><h3>Try Exploratory!</h3><p>If this idea resonates with you, that in the AI era, the ability to <em>read and understand code</em> matters more than ever, then the best way to experience it is to try it yourself.</p><p>In Exploratory, you can describe how you want to transform your data, and AI will generate <strong>readable R (tidyverse) code</strong>for you. Not as a black box, but as something you can inspect, understand, and build upon.</p><p>You’ll quickly notice the difference.</p><p>Instead of repeatedly asking AI to regenerate answers, you begin to <strong>work with the analysis itself</strong>, verifying it, refining it, and evolving it over time.</p><p>👉 <strong>Download Exploratory</strong></p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an account yet, you can sign up here and start a <strong>30-day free trial</strong>:</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>If your trial has expired but you’d like to try the latest AI features, simply launch the newest version and use the <strong>Extend Trial</strong> option.</p><p>If you have any questions or feedback, feel free to reach out directly: <a href="mailto:kan@exploratory.io">kan@exploratory.io</a></p><p>I’d love to hear how you’re using Exploratory to better understand your data.</p><p>Kan Nishida</p><p>CEO, Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=4e9b59beeabd" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/in-the-age-of-ai-rs-readability-becomes-a-superpower-4e9b59beeabd">In the Age of AI, R’s Readability Becomes a Superpower</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Why Deep Learning Didn’t Replace Tree Models for Tabular Data]]></title>
            <link>https://medium.com/learn-dplyr/why-deep-learning-didnt-replace-tree-models-for-tabular-data-d80b796d652f?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/d80b796d652f</guid>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[deep-learning]]></category>
            <category><![CDATA[machine-learning]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Mon, 06 Apr 2026 02:38:19 GMT</pubDate>
            <atom:updated>2026-04-06T02:38:19.186Z</atom:updated>
            <content:encoded><![CDATA[<h4><em>Why models like XGBoost and LightGBM still dominate structured data problems.</em></h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*57av6R0oTPPTCwFdClUrag.png" /></figure><p>Over the past decade, the world of machine learning and AI has been dominated by one idea: deep learning.</p><p>Neural Networks — the algorithms behind deep learning — has transformed fields like:</p><ul><li>computer vision</li><li>speech recognition</li><li>natural language processing</li></ul><p>Neural networks have transformed fields like image recognition, speech processing, and natural language understanding. Models such as transformers and convolutional neural networks now power everything from ChatGPT to self-driving cars.</p><p>Given that success, it seemed inevitable that deep learning would replace traditional machine learning everywhere.</p><p>But something interesting happened.</p><p>In the world of tabular data, the structured datasets used by most businesses, deep learning never completely took over.</p><p>Instead, algorithms like XGBoost and LightGBM continue to dominate many real-world machine learning applications.</p><p>This sounds surprising to some. But once you understand the nature of tabular data and what deep learning is good at (and not good at), it starts to make sense.</p><h3>Most Businesses Data Are Tabular Data</h3><p>Most organizations are not training AI models on billions of images or internet-scale text corpora.</p><p>Instead, they work with data that looks something like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WB7DjS9qZLqoWyYvh8enBA.png" /></figure><p>Columns represent different concepts:</p><ul><li>age</li><li>income</li><li>purchase behavior</li><li>geographic region</li></ul><p>Unlike images or language, there is no inherent structure connecting these variables.</p><p>Images have spatial structure</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*rCGXp8Ef61wutW0I0eZ-bQ.png" /></figure><p>Language has sequential structure</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*u9M53dXZTKxdowwXaOFQBA.png" /></figure><p>But, tabular data is different. It’s simply a collection of variables describing some phenomenon. And it doesn’t have the kind of structure that images and text have.</p><p>Tabular data</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*fvvO4WciUlAy9g8Wep1V-Q.png" /></figure><p>Deep learning thrives when data has hierarchical structure (images, language).</p><p>Think of the modern Transformer, which is one of the most important algorithms in deep learning and AI today. When it builds a model on a given text data, it takes the sequence of text and the relation between words, sentences, etc. into account, and predicts the next word.</p><p>But, tabular data typically does not have such sequence and structure. Each value in a given variable is typically independent from other values.</p><p>And it turned out that this difference matters a lot.</p><h3>Why Tree Based ML Models Work So Well</h3><p>Machine Learning models like XGBoost, LightGBM, Random Forest, etc. are still the most common algorithms among practitioners who are building prediction models with business data, and they all share a common fundamental architecture.</p><p>That is Decision Trees.</p><p>Decision trees approach the problem in a way that feels very natural for this type of data.</p><p>Instead of learning abstract representations, they learn rules.</p><p>For example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QJOru7Oaroj2hPERkLvwVw.png" /></figure><p>Or:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*vi7udNnaiFT94bw11EWK0w.png" /></figure><p>These kinds of conditional rules often show up in real-world datasets, and Decision trees are extremely good at discovering them.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*L3Ks5NLUisgpEho2OkDc6g.png" /></figure><h3>The Rise of XGBoost</h3><p>Around the mid-2010s, one algorithm in the tree family became very popular.</p><p>That algorithm was XGBoost.</p><p>XGBoost implemented gradient boosting in a highly optimized way and quickly became the default choice for many machine learning practitioners working with tabular data.</p><p>It builds trees sequentially, each one correcting the mistakes of the previous model.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*GrFtE50PoiEA8TD-.png" /></figure><p>For several years, it dominated the eyes of data science practitioners who were building prediction models for business data.</p><p>But as datasets grew larger, people began to encounter a new problem.</p><p>Training these models could take a long time.</p><h3>Enter LightGBM</h3><p>In 2017, researchers at Microsoft introduced a new boosting framework called LightGBM.</p><p>The goal wasn’t to reinvent gradient boosting.</p><p>Instead, the idea was to make boosting lighter.</p><p>The word <em>Light</em> in LightGBM refers to being lightweight in computation and memory usage.</p><p>Several clever design decisions helped achieve this:</p><ul><li>trees grow leaf-wise, focusing computation on the most useful splits</li><li>features are converted into histogram bins to reduce split evaluations</li><li>rows with large gradients are prioritized using GOSS</li><li>sparse features are compressed using exclusive feature bundling (EFB)</li></ul><h3>Level-wise vs Leaf-wise growth</h3><p>Level-wise (XGBoost)</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zRsoL1AXHBVfumTg7jhLgw.png" /></figure><p>Leaf-wise (LightGBM)</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*a2xmvGDdJqcJLz8lX7wzag.png" /></figure><p>LightGBM grows trees where the loss decreases most, focusing computation on the most informative parts of the model.</p><p>Together, these ideas dramatically reduce the amount of computation required to train models.</p><p>For people experimenting with machine learning models by trying many features and hyper-parameters, this speed improvement made a huge difference.</p><h3>Why Deep Learning Often Struggles?</h3><p>It’s not that people stopped trying, in fact many researchers have been trying applying neural networks to tabular datasets.</p><p>But very often, Decision tree based boosting models such as XGBoost, LightGBM, still performed better.</p><p>The reason is surprisingly simple.</p><p>Deep learning excels when the data contains rich internal structure.</p><p>Images contain spatial patterns. Language contains grammatical patterns.</p><p>Tabular data usually does not.</p><p>Instead, tabular datasets often contain:</p><ul><li>diverse mix of variables</li><li>engineered features (artificially created extra variables based on original variables)</li><li>sparse categorical encodings (one-hot encoding)</li><li>nonlinear feature interactions</li></ul><p>Decision trees are well suited to discovering these kinds of patterns.</p><h3>What This Means in Practice</h3><p>In many real-world machine learning projects, the workflow often looks like this:</p><ol><li>Build a baseline model.</li><li>Try a tree boosting algorithm.</li><li>Improve the model through feature engineering and parameter tuning.</li></ol><p>And very often, the models that end up performing best are based on boosting algorithms such as XGBoost, LightGBM, etc.</p><p>These algorithms have become reliable workhorses for tabular data problems.</p><h3>Trying These Models Yourself</h3><p>One of the motivations behind building Exploratory was to make various data science tools easier to use, and building prediction models with machine learning models is one of them.</p><p>In Exploratory, you can train tree based models such as:</p><ul><li>Random Forest</li><li>XGBoost</li><li>LightGBM</li></ul><p>directly from an interactive interface.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*4xO-j6CqRxB-6RH1.png" /></figure><p>You can take a look at <a href="https://exploratory.io/note/exploratory/Introduction-to-LightGBM-eAK7zbZ9">this how-to note</a> for more details on how to use LightGBM.</p><p>Instead of writing large amounts of code, you can focus on exploring your data, building features, and comparing models.</p><p>If you work with tabular data such as customer behavior, financial data, operational metrics, etc., it’s definitely worth trying them.</p><h3>A Final Thought</h3><p>As it turned out, the deep learning revolution didn’t eliminate traditional machine learning algorithms.</p><p>Instead, it clarified something important.</p><p>Different types of data require different tools.</p><p>For images and language, as you know, deep learning dominates.</p><p>But, for tabular data, tree based algorithms like XGBoost and LightGBM remain some of the most powerful methods available.</p><p>One of the things I appreciate about boosting algorithms like XGBoost and LightGBM is how practical they are.</p><p>You don’t need massive infrastructure or complicated neural network architectures. You start with your data, build some features (if required), and let the model discover useful patterns.</p><p>In many cases, the results are surprisingly good!</p><p>In Exploratory, you can train models such as Random Forest, XGBoost, and LightGBM directly from an interactive interface and compare their performance on your dataset.</p><p>Sometimes the easiest way to understand the strengths of these models is simply to see how they perform on your own data.</p><h3>Download Exploratory</h3><p>You can start using XGBoost, LightGBM, Random Forest, and other models today in the latest version of Exploratory.</p><p>👉 Download Exploratory</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an account yet, sign up here to start your 30-day free trial.</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>If your trial has expired, simply launch the latest version and use the Extend Trial option.</p><p>If you have questions or feedback, feel free to contact me at <a href="mailto:kan@exploratory.io">kan@exploratory.io .</a></p><p>We’d love to hear how you’re using Exploratory to uncover insights in your data.</p><p>Kan Nishida</p><p>CEO, Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=d80b796d652f" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/why-deep-learning-didnt-replace-tree-models-for-tabular-data-d80b796d652f">Why Deep Learning Didn’t Replace Tree Models for Tabular Data</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[LightGBM Explained: How It Differs from Random Forest and XGBoost]]></title>
            <link>https://medium.com/learn-dplyr/lightgbm-explained-how-it-differs-from-random-forest-and-xgboost-286836838fe7?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/286836838fe7</guid>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[analytics]]></category>
            <category><![CDATA[machine-learning]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Sun, 22 Mar 2026 17:03:57 GMT</pubDate>
            <atom:updated>2026-03-22T17:41:54.095Z</atom:updated>
            <content:encoded><![CDATA[<p><em>The evolution of tree-based models — from robustness to optimization to scalability</em></p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*dhq6nuirJ7cliKbQ.png" /></figure><p>If you work with tabular data (table data), the kind of structured data found in business analytics, finance, marketing, or operations, you’ve probably encountered three popular machine learning algorithms:</p><ul><li>Random Forest</li><li>XGBoost</li><li>LightGBM</li></ul><p>All three rely on decision trees, and they can all produce very strong predictive models. But they were designed with different priorities in mind.</p><p>Random Forest emphasizes simplicity and robustness. XGBoost focuses on highly optimized gradient boosting. LightGBM was created to make boosting faster and more scalable.</p><p>So, which one to choose?</p><p>It depends…</p><p>And this is why I wrote this blog post.</p><p>Understanding why LightGBM was created and how it works makes it much easier to decide whether it’s the right tool for your problem.</p><h3>The Problem LightGBM Was Designed to Solve</h3><p>By the mid-2010s, gradient boosting had already proven to be one of the most powerful techniques for predictive modeling on tabular data, or structured data if you will.</p><p>In particular, XGBoost had become extremely popular after dominating many machine learning competitions.</p><p>However, as datasets continued to grow, practitioners began to encounter new challenges.</p><p>The performance.</p><p>Datasets with millions of rows become normal, and variables (or features) grow thousands, which caused sparse features produced by one-hot encoding.</p><p>Gradient boosting was powerful, but it could also be computationally heavy. This means that it takes time to build models with boosting algorithms when the data size is big.</p><p>Researchers at Microsoft set out to redesign parts of the algorithm so it could handle large datasets more efficiently.</p><p>The result was LightGBM, and they released it as open source in 2017.</p><h3>Why Is It Called “LightGBM”?</h3><p>LightGBM stands for ‘Light Gradient Boosting Machine’.</p><p>The word “Light” does not refer to the speed of light.</p><p>Instead, it refers to the algorithm being lightweight in computation and memory usage.</p><p>The goal was to create a gradient boosting system that could:</p><ul><li>train faster</li><li>use less memory</li><li>scale to larger datasets</li></ul><p>while still maintaining strong predictive performance.</p><p>Before diving into LightGBM’s innovations, it helps to understand how the three algorithms differ conceptually.</p><h3>Random Forest: Many Independent Trees</h3><p>Random Forest builds many trees independently.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*89Fmf8FFrhdNfVbpMlQgOw.png" /></figure><p>Each tree:</p><ol><li>samples the dataset randomly</li><li>selects variables (or features) randomly when splitting</li><li>produces its own prediction</li></ol><p>The final prediction is simply the average (regression) or majority vote (classification).</p><p>Key idea is that many independent trees reduce variance and improve stability compared to a single tree (Decision Tree).</p><p>But the trees do not learn from each other.</p><h3>XGBoost: Trees That Correct Mistakes</h3><p>XGBoost uses a technique called gradient boosting. Instead of building independent trees, it builds trees sequentially to correct the errors made by the previous trees.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*fQR2leC8wtZ44U0AoAtbeQ.png" /></figure><p>Boosting algorithms are really performing a form of gradient descent in function space.</p><p>It builds the first tree and predict and calculate the errors (or loss). If it was a regression problem then the errors can be the residual between the actual and predicted values. And this is called ‘Gradient’.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JNaiZyOh6rK_7ipcddk_7w.png" /></figure><p>The next tree will be built to predict the gradient values, not the actual values because the gradient tells the model how to move predictions to reduce loss.</p><p>And combining the predicted value from the first tree and the predicted values from the second tree will become the new predicted values as a model.</p><blockquote>prediction_new = prediction_from_first_tree + learning_rate × prediction_from_second_tree</blockquote><p>Conceptually, the model evolves like this:</p><ul><li>Tree1 → initial prediction</li><li>Tree2 → fix errors from Tree1</li><li>Tree3 → fix remaining errors</li><li>Tree4 → continue improving</li></ul><p>This process usually leads to more accurate models than Random Forest.</p><p>However, as datasets grow larger, training boosting models can become computationally expensive.</p><p>That is where LightGBM comes in.</p><h3>LightGBM: Designed to scale</h3><p>LightGBM does not change the basic idea of gradient boosting.</p><p>Instead of optimizing only the boosting algorithm, LightGBM also optimizes:</p><ul><li>how trees grow</li><li>how rows are sampled</li><li>how splits are evaluated</li><li>how features are represented</li></ul><p>to scale to very large datasets while maintaining strong accuracy.</p><p>These improvements come from four key innovations:</p><ul><li>Leaf-wise tree growth</li><li>Histogram-based splitting</li><li>GOSS (Gradient-based One-Side Sampling)</li><li>EFB (Exclusive Feature Bundling)</li></ul><p>Let’s walk through them one by one.</p><h4>1. Leaf-Wise Tree Growth</h4><p>One of the most distinctive features of LightGBM is leaf-wise tree growth.</p><p>Traditional tree algorithms such as Random Forest and XGBoost grow trees level-wise.</p><p>At each depth of the tree, all nodes are expanded.</p><p>Example:</p><pre>        Root<br>       /    \<br>      A      B<br>     / \    / \<br>    C   D  E   F</pre><p>Every level of the tree expands evenly, and it produces balanced trees, which are stable and predictable.</p><p>But, they may waste computation expanding branches that do not significantly improve predictions.</p><p>LightGBM takes a different approach. Instead of expanding all nodes at the same depth, it expands the leaf that produces the greatest reduction in loss.</p><p>Example:</p><pre>        Root<br>       /    \<br>      A      B<br>     / \<br>    C   D<br>   /<br>  E</pre><p>The tree grows where the model improves most.</p><p>This approach allows LightGBM to reach strong predictive performance with fewer splits.</p><p>The trade-off is that trees can become deeper in certain branches, so LightGBM provides parameters such as max_depth and num_leaves to control model complexity.</p><h3>2. Histogram-Based Splitting</h3><p>Another important optimization in LightGBM is histogram-based splitting.</p><p>Standard tree algorithms may evaluate many possible split thresholds for continuous features.</p><p>Example:</p><pre>Age ≤ 21<br>Age ≤ 22<br>Age ≤ 23<br>Age ≤ 24</pre><p>LightGBM speeds this up using histogram binning.</p><p>Instead of evaluating every unique value, continuous features are grouped into bins.</p><p>Example:</p><p>Original values:</p><pre>23, 25, 27, 29, 35</pre><p>Converted to bins:</p><pre>20–25<br>25–30<br>30–40</pre><p>Now the algorithm evaluates splits only on bin boundaries.</p><p>This dramatically reduces the number of candidate splits and speeds up training.</p><h3>3. GOSS (Gradient-Based One-Side Sampling)</h3><p>Training boosting models on large datasets normally requires processing all rows.</p><p>LightGBM introduces GOSS (Gradient-based One-Side Sampling) to reduce the number of rows used during training.</p><p>The key idea is based on how boosting works.</p><p>In boosting algorithms, each data point has a gradient value indicating how much the model needs to adjust its prediction.</p><ul><li>Rows with large gradients represent predictions where the model is making large errors.</li><li>Rows with small gradients are already well predicted.</li></ul><p>While typical boosting algorithms including XGBoost randomly sample the data, LightGBM uses this gradient information to sample the data.</p><ul><li>Keeps all rows with large gradients</li><li>Keeps only a subset of rows with small gradients</li></ul><p>Example:</p><p>Dataset: 100,000 rows</p><ul><li>Top 20% largest gradients → keep all (20,000 rows)</li><li>Remaining 80% → sample 10% (8,000 rows)</li></ul><p>This wya, it need to use only 28,000 rows instead of 100,000 rows.</p><p>This significantly reduces computation while preserving important learning signals.</p><h3>4. EFB (Exclusive Feature Bundling)</h3><p>Many modern datasets contain high-dimensional sparse features, especially when categorical variables are one-hot encoded.</p><p>Example:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*_l3nJ6rkcW-4OmX2q6lMIw.png" /></figure><p>These features are <strong>mutually exclusive,</strong> only one can be active in each row.</p><p>Instead of treating them separately, LightGBM bundles them into one feature.</p><p>Original:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MROexJoGbDLdL38-bUYclw.png" /></figure><p>Bundled:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*2f1vmdn0k3vqy7WSUG5N0w.png" /></figure><p>Now the algorithm evaluates splits on <strong>one feature instead of three</strong>.</p><p>This reduces feature dimensionality and speeds up training.</p><h3>How to Choose Each Model?</h3><p>Each algorithm has strengths.</p><h4>Random Forest</h4><p>Good when:</p><ul><li>you want a <strong>simple baseline</strong></li><li>datasets are <strong>relatively small</strong></li><li>minimal tuning is preferred</li></ul><h4>XGBoost</h4><p>Good when:</p><ul><li>datasets are <strong>moderate</strong> in size</li><li>you want strong predictive performance</li><li>stability and extensive tuning options are important</li></ul><h4>LightGBM</h4><p>LightGBM works particularly well when:</p><ul><li>datasets are <strong>large</strong></li><li>feature dimension is <strong>high</strong></li><li>features are <strong>sparse</strong></li><li>training time matters</li></ul><h4>Practical Recommendation</h4><p>A common workflow in machine learning projects is:</p><ol><li>Start with <strong>Random Forest</strong> as a baseline.</li><li>Try <strong>XGBoost</strong> or <strong>LightGBM</strong> to improve performance.</li><li>Prefer <strong>LightGBM</strong> when datasets become large or training time becomes a bottleneck.</li></ol><p>Because of its efficiency and scalability, LightGBM has become a popular choice for many machine learning tasks with tabular data.</p><h3>Final Thought</h3><p>Random Forest, XGBoost, and LightGBM all rely on decision trees, but they represent different philosophies:</p><ul><li>Random Forest focuses on <strong>robust ensembles</strong></li><li>XGBoost focuses on <strong>optimized gradient boosting</strong></li><li>LightGBM focuses on <strong>efficient and scalable boosting</strong></li></ul><p>Understanding these differences helps you choose the right tool, and explains why LightGBM has become an important algorithm in modern machine learning.</p><h3>Try LightGBM with Exploratory!</h3><p>You can try LightGBM with <a href="https://exploratory.io/">Exploratory</a> v14.5 or later versions.</p><ul><li>Go to Analytics view.</li><li>Select LightGBM.</li><li>Select a Target Variable.</li><li>Select Explanatory Variables (Features)</li><li>Click Run button.</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*NTjiAbBsm8qzy0if.png" /></figure><p>You can take a look at <a href="https://exploratory.io/note/exploratory/Introduction-to-LightGBM-eAK7zbZ9">this how-to note</a> for more details on how to use LightGBM.</p><h4>Download Exploratory</h4><p>You can start using LightGBM today in the latest version of Exploratory.</p><p>👉 Download Exploratory v14</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an account yet, sign up here to start your 30-day free trial.</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>If your trial has expired but you’d like to try the new features, simply launch the latest version and use the Extend Trial option.</p><p>If you have questions or feedback, feel free to contact me at <a href="mailto:kan@exploratory.io">kan@exploratory.io .</a></p><p>We’d love to hear how you’re using Exploratory to uncover insights in your data.</p><p>Kan Nishida<br>CEO, Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=286836838fe7" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/lightgbm-explained-how-it-differs-from-random-forest-and-xgboost-286836838fe7">LightGBM Explained: How It Differs from Random Forest and XGBoost</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Auto-Positioning Labels: Keep Your Charts Readable Automatically]]></title>
            <link>https://medium.com/learn-dplyr/auto-positioning-labels-keep-your-charts-readable-automatically-3bbd1f1a0ed8?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/3bbd1f1a0ed8</guid>
            <category><![CDATA[exploratory-data-analysis]]></category>
            <category><![CDATA[data-visualization]]></category>
            <category><![CDATA[data-science]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Thu, 19 Mar 2026 11:18:39 GMT</pubDate>
            <atom:updated>2026-03-19T11:18:39.961Z</atom:updated>
            <content:encoded><![CDATA[<h4>Introducing auto-positioning for chart values in Exploratory</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*0ZaCJYxsHlguMW4K.png" /></figure><p>Showing values directly on a chart is incredibly powerful.</p><p>Your audience first sees the pattern in the data. Then they see the exact labels and numbers behind the pattern.</p><p>That combination is often what makes a chart insightful.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*N5WjChl4RnFRC8I7.png" /></figure><p>Most charting systems place labels exactly at the coordinates of the data point.</p><p>For small datasets, this works perfectly.</p><p>But when you have many data points that are close to each other, your chart would become something like this.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*HICPwfUgv-_vBEvD.png" /></figure><p>Instead of helping the reader, the labels start fighting each other.</p><p>Typical problems appear immediately:</p><ul><li>Labels overlap each other</li><li>Labels cover the data points</li><li>Values become unreadable</li><li>Analysts start manually adjusting labels</li></ul><p>What was supposed to clarify the chart ends up making it harder to read.</p><p>This problem is known as label overlap or label collision.</p><p>And it turns out to be surprisingly difficult to solve.</p><h3>The Idea: Let Labels Move (Just a Little)</h3><p>In Exploratory v14.3 we introduced Auto-Positioning for Labels.</p><p>The idea sounds simple:</p><blockquote><em>Instead of fixing labels exactly on the data points (or coordinates), allow them to move slightly until they no longer overlap.</em></blockquote><p>But there are a few important constraints.</p><p>The labels must:</p><ul><li>Avoid overlapping each other</li><li>Avoid being overlapped by marker (e.g. line, bar, etc.)</li><li>Stay visually close to the original point</li><li>Indicate clearly which data point it represents</li><li>Preserve readability</li></ul><p>When the system finds a better layout, the labels reposition automatically.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*yU6XUaOz67Ageq-5.png" /></figure><p>If a label moves away from its original point, a leader line (arrow) connects the label back to the data point.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*MbZujWD2Zy--FoNo.png" /></figure><p>The result is a chart that remains readable even with many labels.</p><h3>Our First Attempt: Physics Simulation</h3><p>This is a solved problem in R with a package called ‘ggrepel’, but while Exploratory is built on top of R system we use Plotly JS (Java Script) for chart rendering. So we needed to build an auto-positioning system in JS layer.</p><p>So, our first approach was to use d3-force, a physics simulation library in JS.</p><p>The idea was appealing.</p><p>Labels repel each other like particles with repulsive force.</p><p>Eventually they should settle into positions where nothing overlaps.</p><p>In theory.</p><p>In practice, it didn’t work.</p><p>When many values are densely packed the labels tend to push each other away and scatter outside the chart area or across the entire screen.</p><p>So we gave up on that approach.</p><h3>The Algorithm That Worked: Simulated Annealing</h3><p>Instead, we adopted an algorithm called Simulated Annealing, which was used in D3-Labeler.</p><p>This is a classic optimization technique inspired by metallurgy.</p><p>When metal cools slowly, its atoms settle into a stable structure.</p><p>Simulated Annealing follows a similar idea:</p><ol><li>Start with the current label layout</li><li>Move labels slightly in random directions</li><li>Evaluate whether the layout improves</li><li>Gradually reduce the randomness over time</li></ol><p>After many small adjustments, the system converges toward a layout with minimal overlaps and good readability.</p><p>The key advantage is that labels stay close to their original positions, rather than flying across the chart.</p><h3>Engineering the System</h3><p>To integrate this into Exploratory, we created a new module that wraps the chart rendering process.</p><p>The workflow looks like this:</p><ol><li>Exploratory renders the chart normally</li><li>The wrapper module analyzes the label positions</li><li>The auto-positioning algorithm runs</li><li>Labels are repositioned if collisions are detected</li></ol><p>This adjustment happens automatically after the chart rendering.</p><h3>Keeping It Fast (Even with Hundreds of Labels)</h3><p>Optimization algorithms can become slow when the number of labels increases.</p><p>To keep performance smooth, we introduced spatial grid partitioning.</p><p>Instead of checking collisions between every pair of labels, the system divides the chart into small grid cells.</p><p>Labels only need to check nearby neighbors inside the same grid.</p><p>This dramatically reduces the number of collision checks.</p><p>As a result, the system remains responsive even with 500+ labels.</p><h3>Additional Improvements</h3><p>We added a few additional features to make the system more practical.</p><h3>Leader Lines</h3><p>If a label moves far from its original point, a leader line automatically appears.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*-HHUEcq1u6kqMrq7.png" /></figure><p>This maintains the visual connection between label and data point.</p><h3>Error Bar Awareness</h3><p>Bar charts with error bars introduce another challenge.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*904GVU9JWcGHAHDV.png" /></figure><p>Labels must avoid overlapping the confidence interval lines.</p><p>The algorithm takes error bar ranges into account when calculating positions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*EjnGqJTRXkuIctIc.png" /></figure><h3>How to Enable Auto-Positioning</h3><p>Using the Auto-Positioning feature is simple.</p><ol><li>Open the chart property dialog.</li><li>Click the gear icon at the top of the chart.</li><li>Then go to the Values tab.</li><li>Enable Show Values.</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*_LZ8KK1ot46MRK3_.png" /></figure><p>When the Position is set to Automatic, Exploratory automatically adjusts label positions to avoid overlaps.</p><p>If labels move far from their original point, arrows appear to indicate the connection.</p><p>You can control the color and the arrow visibility using the Arrow Display Threshold setting.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*OTNOi2WxtvIggLtI.png" /></figure><p>Increasing the threshold hides shorter arrows and reduces clutter.</p><p>For a newly created charts, when you enable to show the values (labels) on chart the auto-positioning is automatically set by default. For existing charts that you created before v14.3, you want to manually switch the Position to Automatic.</p><h3>Improving Placement Accuracy</h3><p>You can also control the optimization effort.</p><p>The setting ‘# Tries to Improve Accuracy’ determines how many optimization trials are performed.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3EDsvFapEzU0ecidjiN-9Q.png" /></figure><ul><li>Increase the value → more accurate placement (slower)</li><li>Decrease the value → faster calculation (less precise)</li></ul><h3>A Small Feature That Makes Powerful Data Exploration</h3><p>In Exploratory v14.3 we introduced this auto-positioning system using Simulated Annealing and spatial optimization.</p><p>At first glance, this might look like a small feature.</p><p>But in practice, it changes something fundamental about how you explore data.</p><p>Exploratory Data Analysis is not just about creating charts.</p><p>It is about discovering things you did not expect to see.</p><p>That kind of discovery often happens when you start looking closely at individual data points.</p><p>For example, imagine a scatterplot showing 190 countries.</p><p>At first, you might look for a particular country you are interested in.</p><p>But once the labels become readable, something else begins to happen.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*PDFz0VpltJYt2OQN.png" /></figure><p>You start to notice patterns you did not plan to look for.</p><ul><li>Which countries are close to each other?</li><li>Which countries behave similarly?</li><li>Which countries stand apart from the rest?</li></ul><p>These moments of discovery are the essence of exploratory data analysis.</p><p>However, if the labels overlap and become unreadable, those discoveries become much harder.</p><p>Analysts end up spending time manually adjusting labels instead of exploring the data itself.</p><p>Our goal with Exploratory has always been to build a tool that helps people think better with data.</p><p>Not by automating the thinking process, but by removing the friction that gets in the way of exploration.</p><p>Auto-positioning labels is one of those small features that quietly makes exploration easier.</p><p>And when exploration becomes easier, new insights often follow.</p><p>That is why we believe this feature helps make Exploratory a better environment for true Exploratory Data Analysis.</p><h3>Try Auto-Positioning Today</h3><p>You can start using the Auto-Positioning feature today in the latest version of Exploratory.</p><p>👉 Download Exploratory v14</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an account yet, sign up here to start your 30-day free trial.</p><p><a href="https://exploratory.io/">https://exploratory.io/</a></p><p>If your trial has expired but you’d like to try the new features, simply launch the latest version and use the Extend Trial option.</p><p>If you have questions or feedback, feel free to contact me at <a href="mailto:kan@exploratory.io">kan@exploratory.io .</a></p><p>We’d love to hear how you’re using Exploratory to uncover insights in your data.</p><p>Kan Nishida</p><p>CEO, Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=3bbd1f1a0ed8" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/auto-positioning-labels-keep-your-charts-readable-automatically-3bbd1f1a0ed8">Auto-Positioning Labels: Keep Your Charts Readable Automatically</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[AI Note Editor: Create Reports 10x Faster, 10x Better]]></title>
            <link>https://medium.com/learn-dplyr/ai-note-editor-create-reports-10x-faster-10x-better-8249816ffe48?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/8249816ffe48</guid>
            <category><![CDATA[ai]]></category>
            <category><![CDATA[data-scien]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Mon, 16 Mar 2026 11:56:12 GMT</pubDate>
            <atom:updated>2026-03-16T11:56:26.211Z</atom:updated>
            <content:encoded><![CDATA[<p>We’re thrilled to introduce <strong>AI Note Editor</strong> in Exploratory v14! 🎉</p><p>This new feature works like having a professional editor — and a data analyst — right beside you as you write. Except it never gets tired, always respond fast, and can analyze your charts with deep statistical knowledge.</p><p>It’s designed to <strong>help you create high-quality analysis reports quickly, clearly, and effortlessly.</strong></p><h3>Why Reporting Is the Most Underrated (and Most Painful) Step in Data Analysis</h3><p>In any real-world data analysis workflow, running the analysis is only half the job. The <em>other</em> half — often the harder half — is explaining what you discovered.</p><p>But let’s be honest: most of us <em>don’t</em> enjoy writing reports. Sometimes, we aren’t sure <strong>how to describe what charts are showing</strong>. So, we end up sending a Slack message with “here’s the chart” or copying &amp; pasting charts in PowerPoint/Slides and hoping the audience magically understands it.</p><p>As a result:</p><ul><li>Screenshots pile up in Slack threads.</li><li>PowerPoint slides become chart image dumps.</li><li>Insights get lost because they’re never clearly communicated.</li></ul><p>This is the communication problem in Data Science workflow that we wanted to solve with the new AI Note Editor.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*5kGLUn6-k_cub-SILa-o5Q.png" /></figure><h3>Meet AI Note Editor: Turn Comments &amp; Charts Into a Clear, Polished Report</h3><p>With AI Note Editor, your analysis report writing workflow will become something like this.</p><ol><li>Add your charts.</li><li>Write a few comments, if you like.</li><li>Get a complete data analysis reports generated.</li><li>Edit it as it fits your needs.</li><li>Share your report with others!</li></ol><p>Based on your comments and charts, AI Note Editor generates a <strong>polished, structured, ready-to-share report</strong> for you, right inside Exploratory.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WUoySat756_1AibQWIaobA.png" /></figure><p>No switching tools.</p><p>No copy &amp; paste to Power-Point.</p><p>No asking somebody else to create reports.</p><p>No staring at a blinking cursor trying to think of the right words.</p><p>Your data analysis stays where it belongs — inside the Exploratory workflow — and AI handles the writing.</p><h3>Automatic Chart Interpretation</h3><p>AI Note Editor doesn’t just improve your writing, it can also <strong>interpret your charts</strong> and explain what’s happening in the data. This is something that no general-purpose writing tool can do.</p><p>Give it a chart and the AI will:</p><ul><li>Identify key patterns</li><li>Explain strengths and weaknesses</li><li>Describe trends, anomalies, or outliers</li><li>Summarize relationships between variables</li><li>Highlight important signals</li><li>Put everything into intuitive, natural language</li></ul><p>For example, given a radar chart like the one below, AI will describe the key patterns, strengths, weaknesses, and overall story behind the data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ZUDUoBlwjresAk7ziyV8-w.png" /></figure><h4>Detect Trends &amp; Signals</h4><p>Chart interpretation not only interprets numerical values but also provides context-aware interpretations, such as trends and signals within the data.</p><p>For example, with XmR charts (control charts), it can detect whether a signal is present and explain what it indicates accordingly.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*OhfeHGHZoHT1fSdS3e2-1w.png" /></figure><h4>AI Note Editor that comes with a deep statistical knowledge</h4><p>Moreover, for charts containing statistical information, the AI Note Editor provides interpretations based on that statistical data.</p><p>For example, for a scatter plot showing the relationship between two variables, it can explain how strong (or weak) the correlation is, what the relationship implies, and whether it is statistically significant or not.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*Ph3oxT0qeSRo3eKwOHgMBQ.png" /></figure><h4>AI Note Editor Can Analyze Data and Explain What’s There</h4><p>You’ve probably experienced something like these before:</p><ul><li>“I know what I see with the chart… but how do I put it into words?”</li><li>“I think this trend is important, but I’m not confident how to explain it.”</li><li>“I saw the chart, but I didn’t even notice it until someone pointed it out.”</li></ul><p>AI Note Editor can address such concerns by:</p><ul><li>Helping you verbalize what the data shows</li><li>Pointing out things you might have missed</li><li>Ensuring your analysis is complete</li><li>Reducing the risk of overlooking important signals</li><li>Improving the clarity and quality of your reports — instantly</li></ul><p>AI Note Editor is not just about writing faster. It’s about <strong>analyzing data better</strong> and <strong>communicating more clearly</strong>.</p><h3>It can help you Write Better</h3><p>AI Note Editor also includes a set of tools to improve your writing:</p><ul><li>Summarize long text</li><li>Fix grammar and spelling</li><li>Improve clarity and tone</li><li>Refine wording and expression</li><li>Translate your writing</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*CbF-TnbIT3vk9iNp2a0S_A.png" /></figure><p>Whether you’re writing an internal update, a weekly KPI brief, or a full analysis report, AI Note Editor makes the report writing process faster and better.</p><h3>Create Your Own Report Format with Custom Prompts</h3><p>Exploratory provides an “Analysis Report” style out of the box — but let’s be honest, <strong>not everyone writes reports the same way.</strong></p><p>You have your own reporting needs and might need to write in different formats or style to fit your clients or audience’s needs by project to project.</p><p>That’s why AI Note Editor includes <strong>Custom Prompts </strong>support.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ZhxkebrAIf3BXzrml-eh-w.png" /></figure><p>Just use ‘<strong>Run Custom Prompt</strong>’ and describe the format you want with Markdown instructing a specific structure with:</p><ul><li>Headings and subheadings</li><li>Bullet points</li><li>Summary sections</li><li>Business-style executive summaries</li><li>Step-by-step analysis</li><li>Narrative storytelling</li><li>Even templates that match your corporate writing guidelines</li></ul><p>Just tell the AI the structure you want — and it will generate the full report in that format.</p><h3>Save &amp; Reuse Prompt Templates</h3><p>Once you create a prompt you like, you can save it as a <strong>template</strong>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*mvJid8LtKTTDu43hyH1pMQ.png" /></figure><p>Your templates appear in the Template list, ready to use anytime.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QZdBW9EtvGYSA6xAjE_IKg.png" /></figure><p>This is incredibly useful for reports you produce regularly — weekly reports, monthly summaries, recurring analyses, etc. Just pick a template and generate a polished report in seconds — consistent, clean, and always in the right style.</p><h3>A Gallery of Prompt Examples to Get You Started</h3><p>To help you get started, we have prepared a collection of prompt examples in the <a href="https://exploratory.io/tag/?sort=&amp;language=ja&amp;q=tag%3A%22Ai%20Note%20Editor%22&amp;searchType=keyword">AI Note Editor Gallery</a>. You can browse through the prompts, copy &amp; paste them, and tweak however you like.</p><p>You can create sophisticated report formats immediately — even if you’ve never written a prompt before.</p><h3>Share Templates with Your Team</h3><p>AI Note Editor templates can also be <strong>shared across teams</strong>.</p><p>Your team members can directly import them by clicking the “Import” button from the “AI Prompt Template” dialog mentioned above.</p><figure><img alt="" src="https://cdn-images-1.medium.com/proxy/1*MDVv4nYpEzl8fOlolhcvMA.png" /></figure><p>This means:</p><ul><li>Standardized reporting</li><li>Consistent communication style</li><li>Faster onboarding for new members</li><li>Reduced back-and-forth editing</li><li>Higher-quality reports across the organization</li></ul><h3>A New Standard for Data Communication</h3><p>AI Note Editor isn’t just a writing tool — it’s a communication tool with a deep statistical knowledge. It helps you:</p><ul><li>Turn analysis into narrative</li><li>Turn charts into discoveries and insights</li><li>Turn discoveries into stories</li><li>Turn insights into action</li></ul><p>No more copy &amp; paste dumps in Power-Point slides.</p><p>No more staring at charts wondering how to describe what you see.</p><p>No more wasting time wondering what to write and how to explain.</p><p>AI Note Editor gives you clarity.</p><p>Your audience gets deeper understanding of your discoveries.</p><p>And your analysis becomes dramatically more impactful.</p><p>This is what the future of data communication looks like — and it lives directly inside Exploratory.</p><h3>Try AI Note Editor Today!</h3><p>You can start using AI Note Editor right now in the latest version of Exploratory.</p><p>👉 <strong>Download Exploratory v14</strong><br> <a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an Exploratory account yet, please <a href="https://exploratory.io/">sign up here</a> to try it out. The first 30 days are a free trial period!</p><p>If your trial has already expired but you want to try the new AI features, simply launch the latest version and use the “Extend Trial” option in the dialog — or contact us directly.</p><p>For questions or feedback, feel free to reach out:<br> 📧 <strong>support@exploratory.io</strong></p><p>We can’t wait to see the reports you create — and how AI Note Editor helps you communicate insights faster, clearer, and more confidently than ever before.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=8249816ffe48" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/ai-note-editor-create-reports-10x-faster-10x-better-8249816ffe48">AI Note Editor: Create Reports 10x Faster, 10x Better</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Turn GitHub Issues into Release Notes with AI]]></title>
            <link>https://medium.com/learn-dplyr/turn-github-issues-into-release-notes-with-ai-8857dc7f1f27?source=rss-1bfa80768afa------2</link>
            <guid isPermaLink="false">https://medium.com/p/8857dc7f1f27</guid>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[ai]]></category>
            <dc:creator><![CDATA[Kan Nishida]]></dc:creator>
            <pubDate>Mon, 16 Mar 2026 11:42:27 GMT</pubDate>
            <atom:updated>2026-03-16T12:01:54.398Z</atom:updated>
            <content:encoded><![CDATA[<h4>How I Use Exploratory’s AI Function to Automatically Categorize, Rewrite, and Generate Release Notes</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zi3ttrkn75iE1EYPw7gVhw.png" /></figure><p>Preparing a release note used to take a few hours every release. With AI Function, it now takes less than a few minute.</p><p>Every time we ship a new version, we go through dozens of GitHub issues — bug fixes, enhancements, and new features — and turn them into a release note that our users can easily understand.</p><p>This used to be a very manual process.</p><p>For each issue we had to:</p><ol><li>Read the issue title and description</li><li>Decide whether it is a bug fix, enhancement, or documentation change</li><li>Assign it to a functional category (Data Wrangling, Chart, Analytics, etc.)</li><li>Rewrite the title so users understand what changed and why it matters</li></ol><p>When there are 50–100 issues per release, this quickly becomes tedious.</p><p>Today, we automate most of this workflow using AI Function and AI Note Editor inside Exploratory.</p><p>In this post, I’ll show you exactly how we do it.</p><p>My hope is that this gives you ideas for how you can use AI Function to automate your own text-based workflows.</p><h3>What is AI Function?</h3><p>AI Function in Exploratory lets you create a function using a prompt so that AI (LLM) processes each row of your data.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*wUFP3AAbAeKx1clP.png" /></figure><p>You can ask AI to:</p><ul><li>analyze text</li><li>classify information</li><li>summarize content</li><li>generate new text</li><li>transform messy data</li></ul><p>all directly inside your data workflow.</p><p>For example, if you have a column containing customer feedback comments, you can simply write a prompt like:</p><blockquote>Classify the text into several groups.</blockquote><p>AI will analyze each row and assign a category.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-J-Gk6nmRGGnFFnlNOxBYw.png" /></figure><p>You can learn more about AI Function here.</p><ul><li><a href="https://blog.exploratory.io/data-science-2-0-a-new-era-of-text-data-analysis-b6430baadba7">https://blog.exploratory.io/data-science-2-0-a-new-era-of-text-data-analysis-b6430baadba7</a></li></ul><p>I personally use AI Function almost every day.</p><p>When we built this feature, I realized something interesting:</p><blockquote><em>There was far more text data in my daily work than I had ever noticed before.</em></blockquote><p>Issue logs.</p><p>Customer feedback.</p><p>Meeting notes.</p><p>Support conversations.</p><p>Task descriptions.</p><p>These are all valuable data sources, but without the right tools they’re difficult to analyze or automate.</p><p>AI Function changes that.</p><h3>Example: Automating Our Release Notes</h3><p>Let me show you a real example from our workflow.</p><p>We manage all development work in GitHub Issues.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*13wKRqHKF41cAQl5G30g6g.png" /></figure><p>These issues include:</p><ul><li>bug reports</li><li>feature requests</li><li>internal development tasks</li><li>documentation updates</li></ul><p>When we release a new version, we publish the closed issues for that milestone as the release note.</p><p>But raw GitHub issue titles are not written for users — they’re written for developers.</p><p>So we need to transform them.</p><p>Specifically, we need to:</p><ol><li>Categorize issues by product area</li><li>Identify whether they are bug fixes or enhancements</li><li>Rewrite the title so users clearly understand the change</li></ol><p>Before AI Function, I did this manually.</p><p>Now it’s automated.</p><h4>Step 1: Import GitHub Issues</h4><p>First, we import GitHub issues directly into Exploratory.</p><p>For example, we can import issues for the milestone v14.5.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ErhMO9DYUfwzfFYWpDudcw.png" /></figure><p>Once imported, the dataset contains columns like:</p><ul><li>Issue title</li><li>Issue body</li><li>Labels</li><li>Status</li><li>Milestone</li></ul><p>Take a look at this how-to note for details.</p><ul><li><a href="https://exploratory.io/note/exploratory/How-to-Install-Github-Issue-Data-uGz1DrV2">How to Import Github Issue Data</a></li></ul><p>Now that the data is imported, it’s time to work with AI Function.</p><h4>Step 2: Categorize Issues by Product Area</h4><p>First, we categorize each issue into a functional area.</p><p>For this, I create an AI Function with a prompt like this:</p><pre>Based on the title and the body text, categorize text to one of the following groups.</pre><pre>AI Function<br>AI Prompt<br>Summary View<br>Table View<br>Data Source<br>Data Wrangling<br>Chart<br>Analytics<br>Note<br>Dashboard<br>Parameter<br>Project<br>Publish<br>Install<br>Document<br>Others</pre><p>Instead of letting AI invent categories, I give it a predefined list.</p><p>This improves consistency.</p><p>I also provide two columns as input:</p><ul><li>Title</li><li>Body</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JwFPIYrrganFeodUIsyEbA.png" /></figure><p>The title text and the body text would look like the below on the actual Github issue page.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WNs1bgTFhr-FHIUdj-1r3Q.png" /></figure><p>This gives the model enough context</p><p>When executed, the AI analyzes the issue text and assigns an appropriate category for each row.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*NNnb7mqyUudAsW2nsg9dhA.png" /></figure><h4>Step 3: Identify Issue Type</h4><p>Next, we classify whether the issue is:</p><ul><li>a bug fix</li><li>an enhancement</li><li>documentation</li><li>other</li></ul><p>The prompt is simple:</p><p>Identify if a given sentence indicates whether it is an issue fix, a product enhancement, documents, or others.</p><blockquote>Identify if a given sentence indicates whether it is an issue fix, a product enhancement, documents, or others.</blockquote><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*YR9CFqSiiC9dQ6jwVbCLcQ.png" /></figure><p>This correctly categorize them into the right group.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*BVn8iI2346LaqxaSc0yinA.png" /></figure><h3>Step 4: Rewrite the Issue Title</h3><p>The final step is improving the issue titles.</p><p>GitHub issue titles are often short and technical.</p><p>For example:</p><blockquote>AI Function: Don’t run the cached step when duplicating the data frame</blockquote><p>That’s clear to developers, but not to most users.</p><p>We already built an internal system to improve and clean up the issue title and generate the Release Note, but we’ve realized that we can simply use AI Function to rewrite a better title for each issue based on the title and the body text</p><p>So I use AI Function again to generate a better description.</p><p>Prompt:</p><blockquote>Based on the ‘title’ and ‘body’ text, write a one or two sentence summary that clearly explains what the issue is and why it matters to users.</blockquote><p>In the AI Function dialog, I’m setting ‘title’ and ‘body’ columns as the Target Columns.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3LRjlRyM0rCdXHX58vByIw.png" /></figure><p>The result:</p><pre>Original: </pre><pre>AI Function: Don&#39;t run the cached step when duplicating the data frame</pre><pre>New:</pre><pre>Duplicating a data frame currently triggers re-execution of AI Function steps, which can cause unnecessary processing time and API costs.</pre><p>Much clearer.</p><p>And this happens for every issue automatically.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3xRb6EdX4a5W7dE_QB_1Cw.png" /></figure><p>That’s all!</p><p>Now that I have categorized the issues and improved the text I can show them in a table format.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ojP3Gi0z3KsrW-NrOBRwrA.png" /></figure><h4>From Raw Issues to Clean Release Notes</h4><p>After these steps, we now have structured information:</p><ul><li>Issue category</li><li>Issue type</li><li>Improved description</li></ul><p>This makes it easy to organize them into a release note.</p><h3>Next Step: Generate Release Note with AI Note Editor</h3><p>Once the data is prepared, we take it one step further.</p><p>We use AI Note Editor inside Exploratory to generate the release note itself.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*o5InimZ9fxVKdpN5IFcAmA.png" /></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*JUpW2nTC-369t7O61s7F7w.png" /></figure><p>Check out this introductory blog post for AI Note Editor.</p><ul><li><a href="https://blog.exploratory.io/ai-note-editor-create-reports-10x-faster-10x-better-8249816ffe48">AI Note Editor: Create Reports 10x Faster, 10x Better</a></li></ul><p>I’ll write another blog post explaining exactly how we do this, so stay tuned!</p><h3>The Real Power: Reproducibility</h3><p>The biggest benefit of this workflow is <strong>reproducibility</strong>.</p><p>Once the AI Function steps are created, we can reuse them.</p><p>What if the new issues have been added in the last minutes?</p><p>You can:</p><ol><li>Re-import issues for the new milestone</li><li>Run the workflow</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*o9PUnOE24MDLnHRmmX1A9Q.png" /></figure><p>What if we will need to generate a release note for another version in future?</p><ol><li>Click on the Data Source step and open the Import dialog</li><li>Update the milestone</li><li>Run the same workflow</li></ol><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qZbM0sBB5TEgEENMqiM8LQ.png" /></figure><p>We release new versions often, so being able to repeat the same workflow automatically is important for us.</p><p>In any future releases I can simply click a button to re-import data, and the AI Functions automatically process the new issues and produce clean, user-friendly descriptions.</p><p>No manual rewriting required.</p><p>This is the real power of combining:</p><ul><li>data workflows</li><li>AI functions</li><li>reproducibility</li></ul><p>That is the real power of building an automated data wrangling system with AI Function that is reproducible and can be used with any future incoming data.</p><h3>Why Not Just Copy and Paste the Issues into ChatGPT?</h3><p>At this point, you might be thinking:</p><blockquote><em>“Couldn’t I just copy the list of issues into ChatGPT and ask it to do the same thing?”</em></blockquote><p>Yes, you certainly could.</p><p>But once you try to use that approach in a real workflow, several practical limitations quickly appear.</p><p>What makes AI Function inside Exploratory powerful is not just that it uses AI. It’s that AI becomes part of a data workflow.</p><p>Here are some key differences.</p><h4>1. Iterative Workflow Development</h4><p>When working with AI Function, you can build your workflow step by step.</p><p>For example:</p><ol><li>Start with issue categorization</li><li>Review the results</li><li>Improve the prompt</li><li>Run it again</li><li>Add another AI Function step for issue type classification</li><li>Add another step to rewrite titles</li></ol><p>Each step produces a visible column of results, which makes it easy to:</p><ul><li>review</li><li>refine prompts</li><li>improve output quality</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*nw4I4WtoYAGqDTsnTfLIxQ.png" /></figure><p>This iterative workflow is much harder to do when you are simply pasting text into an AI chat interface.</p><h4>2. Control Over the Output</h4><p>When AI is part of your data workflow, you have much more control over the result.</p><p>For example, you can:</p><ul><li>specify allowed categories</li><li>combine multiple columns as context</li><li>inspect the output row by row</li><li>filter and fix problematic cases</li></ul><p>This is important if you care about quality and consistency, not just quick answers.</p><h4>3. Reproducibility</h4><p>This is perhaps the most important difference.</p><p>Once you build the workflow, it becomes reproducible.</p><p>Every time you import a new set of issues, you can run the exact same steps and get consistent results.</p><p>With copy-paste AI workflows, you would have to:</p><ul><li>prepare the text</li><li>paste it into the AI</li><li>rewrite prompts</li><li>manually format the results</li></ul><p>every single time.</p><p>AI Function turns that process into something you can run repeatedly with one click.</p><h4>4. AI Works Directly on Your Data</h4><p>Instead of copying and pasting text back and forth, AI Function works directly on the data frame in Exploratory.</p><p>This means you can combine AI with other data operations, such as:</p><ul><li>filtering rows</li><li>joining datasets</li><li>grouping and summarizing results</li><li>visualizing patterns</li></ul><p>AI becomes just another data transformation step inside the workflow.</p><h4>5. Scales to Large Datasets</h4><p>Chat interfaces are great for small tasks, but they quickly become cumbersome when working with hundreds or thousands of rows of data.</p><p>AI Function processes your data row by row, allowing you to apply the same logic consistently across the entire dataset.</p><h4>6. Part of a Larger Data System</h4><p>Finally, AI Function integrates with the rest of Exploratory’s capabilities:</p><ul><li>data wrangling</li><li>visualization</li><li>machine learning</li><li>dashboards</li><li>notes</li></ul><p>In our case, the output of AI Function becomes the input for AI Note Editor, which then generates the release notes themselves.</p><p>This creates a complete pipeline:</p><p>GitHub Issues → AI Function → Enriched Data → AI Note Editor → Release Note</p><h4>The Real Difference</h4><p>Using AI in a chat interface is great for one-off tasks.</p><p>But AI Function allows you to turn those tasks into reusable data workflows.</p><p>And once that happens, something interesting occurs:</p><blockquote><em>AI stops being a tool you occasionally use, and becomes part of your everyday data process.</em></blockquote><p>That’s the real power of AI Function inside Exploratory.</p><h3>Key Takeaways: Why AI Function Is a Game Changer</h3><p>This example is not really about release notes.</p><p>It’s about something much bigger.</p><p>Many everyday workflows involve unstructured text data:</p><ul><li>GitHub issues</li><li>customer feedback</li><li>support tickets</li><li>meeting notes</li><li>task descriptions</li><li>survey comments</li></ul><p>Traditionally, these workflows required manual reading and interpretation, which made them difficult to automate.</p><p>That’s why many teams simply accept them as manual work.</p><p>But AI Function changes this.</p><p>Instead of writing complicated scripts or building custom AI pipelines, you can simply describe the task in plain language and apply it directly to your data.</p><p>In the release note example above, AI Function helped automate tasks that previously required manual effort:</p><ul><li>Categorizing issues by product area</li><li>Identifying whether they are bug fixes or enhancements</li><li>Rewriting issue titles into user-friendly descriptions</li></ul><p>Once the workflow is built, it becomes reusable and reproducible.</p><p>Every new release can go through the exact same process automatically.</p><p>What used to be a repetitive manual task becomes a data workflow you run with one click.</p><p>And this idea applies far beyond release notes.</p><p>Anywhere you have rows of text data, AI Function can help you:</p><ul><li>classify</li><li>summarize</li><li>transform</li><li>generate structured information</li></ul><p>directly inside your data workflow.</p><p>This is why we built AI Function inside Exploratory.</p><p>Not to replace human thinking, but to remove the repetitive parts of working with text, so you can focus on the insights and decisions that matter.</p><p>AI Function turns messy text data into something you can actually work with, and once that happens, many workflows that used to be manual suddenly become automated.</p><h3>Try AI Function with Your Own Data</h3><p>If you work with text data such as:</p><ul><li>issue logs</li><li>support tickets</li><li>customer feedback</li><li>meeting notes</li><li>task lists</li></ul><p>You can build similar workflows with AI Function.</p><p>If you’re new to AI Function, I recommend starting with the examples available in the Create AI Function menu.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*TdwSFHL_kqKANImJSpVmBQ.png" /></figure><h3>Try AI Functions Today!</h3><p>You can start using AI Functions today in the latest version of Exploratory.</p><p>👉 Download Exploratory v14</p><p><a href="https://exploratory.io/download">https://exploratory.io/download</a></p><p>If you don’t have an Exploratory account yet, please <a href="https://exploratory.io/">sign up here</a> to try it out. The first 30 days are a free trial period!</p><p>If your trial has already expired but you want to try the new AI features, simply launch the latest version and use the “Extend Trial” option in the dialog, or contact us (support@exploratory.io) directly.</p><p>If you have any questions or feedback, please contact me at <a href="mailto:kan@exploratory.io">kan@exploratory.io</a></p><p>We’d love to hear what data wrangling system you build with AI Functions!</p><p>Kan</p><p>CEO/Exploratory</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=8857dc7f1f27" width="1" height="1" alt=""><hr><p><a href="https://medium.com/learn-dplyr/turn-github-issues-into-release-notes-with-ai-8857dc7f1f27">Turn GitHub Issues into Release Notes with AI</a> was originally published in <a href="https://medium.com/learn-dplyr">learn data science</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
    </channel>
</rss>