<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Data Analysis]]></title><description><![CDATA[Data Analysis]]></description><link>https://sonuhere.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6a93dde20fe06c4affc21105/6982f6b1-3694-4ab0-bbb7-1472201d863c.jpg</url><title>Data Analysis</title><link>https://sonuhere.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Tue, 29 Sep 2026 10:01:04 GMT</lastBuildDate><atom:link href="https://sonuhere.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Pandas Data Analysis Project]]></title><description><![CDATA[Introduction
This project was created to practice a complete data-analysis workflow using Python and Pandas.
I worked with a financial fraud detection dataset, focusing on data manipulation, structure]]></description><link>https://sonuhere.hashnode.dev/pandas-data-analysis-project</link><guid isPermaLink="true">https://sonuhere.hashnode.dev/pandas-data-analysis-project</guid><category><![CDATA[Python]]></category><category><![CDATA[data analysis]]></category><category><![CDATA[pandas]]></category><dc:creator><![CDATA[Sonu Gupta]]></dc:creator><pubDate>Sun, 30 Aug 2026 09:12:37 GMT</pubDate><content:encoded><![CDATA[<h2>Introduction</h2>
<p>This project was created to practice a complete data-analysis workflow using <strong>Python and Pandas</strong>.</p>
<p>I worked with a <strong>financial fraud detection dataset</strong>, focusing on data manipulation, structured processing, and exploring patterns within the data.</p>
<h2>What I Did</h2>
<p>Using Pandas and Python, I:</p>
<ul>
<li><p>Processed and explored the dataset</p>
</li>
<li><p>Performed data manipulation and transformation</p>
</li>
<li><p>Analyzed patterns within the data</p>
</li>
<li><p>Used Python-based visualization to better understand the results</p>
</li>
</ul>
<h2>Tech Stack</h2>
<p><strong>Python · Pandas · Data Analysis</strong></p>
<h2>What I Learned</h2>
<p>This project helped me strengthen my understanding of <strong>Pandas, dataset exploration, data transformation, and analytical workflows</strong>.</p>
<p>It also gave me more practice working with structured data before moving toward larger analytics projects.</p>
<h2>Project</h2>
<p><strong>GitHub:</strong> <a href="https://github.com/mahii-17/pandas%5C_project">https://github.com/mahii-17/pandas\_project</a></p>
]]></content:encoded></item><item><title><![CDATA[YouTube Analytics Dashboard: Turning Video Data into Actionable Insights]]></title><description><![CDATA[How I built an interactive dashboard using Python, Pandas, NumPy, Plotly, and Streamlit to understand video performance.
Introduction
Creators generate a huge amount of data every time they publish a ]]></description><link>https://sonuhere.hashnode.dev/youtube-analytics-dashboard-turning-video-data-into-actionable-insights</link><guid isPermaLink="true">https://sonuhere.hashnode.dev/youtube-analytics-dashboard-turning-video-data-into-actionable-insights</guid><dc:creator><![CDATA[Sonu Gupta]]></dc:creator><pubDate>Sun, 30 Aug 2026 09:04:33 GMT</pubDate><content:encoded><![CDATA[<p>How I built an interactive dashboard using Python, Pandas, NumPy, Plotly, and Streamlit to understand video performance.</p>
<h2>Introduction</h2>
<p>Creators generate a huge amount of data every time they publish a video.</p>
<p>Views, likes, comments, watch time, engagement, and other performance indicators can all tell a story. But looking at these numbers individually doesn't always make it easy to understand what is actually happening.</p>
<p>For this project, I built a <strong>YouTube Analytics Dashboard</strong> using Python to turn video-performance data into an interactive and easier-to-understand analytical experience.</p>
<p>The dashboard tracks <strong>7+ performance KPIs</strong>, analyzes <strong>30-day video performance</strong>, and uses <strong>percentile-based comparisons</strong> to help evaluate how an individual video performs relative to typical channel performance.</p>
<hr />
<h2>The Problem</h2>
<p>YouTube provides a lot of information about content performance.</p>
<p>The challenge is not simply having access to the numbers.</p>
<p>The real challenge is answering questions such as:</p>
<ul>
<li><p>How is a video performing?</p>
</li>
<li><p>Which performance metrics deserve attention?</p>
</li>
<li><p>Is a video's performance typical for the channel?</p>
</li>
<li><p>How can multiple performance indicators be viewed together?</p>
</li>
<li><p>Can raw analytics data be turned into something easier to interpret?</p>
</li>
</ul>
<p>I wanted to build a dashboard that could bring these pieces together in one place.<br />Project Goal</p>
<p>The main goal of this project was to create an interactive dashboard that could:</p>
<p><strong>Track important YouTube performance metrics</strong></p>
<p><strong>Analyze recent video performance</strong></p>
<p><strong>Compare individual videos with typical channel performance</strong></p>
<p><strong>Present analytical results through interactive visualizations</strong></p>
<p>Instead of looking at raw rows of data, the dashboard provides a more structured way to explore video performance.</p>
<h2>Dataset &amp; Analysis</h2>
<p>The project works with YouTube video-performance data and focuses on analyzing recent performance over a <strong>30-day period</strong>.</p>
<p>The analysis combines multiple metrics so that performance can be viewed from different perspectives rather than relying on a single number.</p>
<p>This is particularly useful because a video with high views does not necessarily tell the complete story of its performance.</p>
<h2>1. Data Preparation with Pandas</h2>
<p>I used <strong>Pandas</strong> to work with the dataset and prepare the information for analysis.</p>
<p>This involved processing the available video data and structuring it in a way that could be used effectively by the dashboard.</p>
<p>Pandas was also used as part of the feature-engineering and analytical workflow.  </p>
<h2>2. Feature Engineering</h2>
<p>Raw metrics are useful, but derived metrics can provide additional analytical context.</p>
<p>For this project, I used feature engineering to create a more meaningful way of analyzing video performance.</p>
<p>This allowed the dashboard to move beyond simply displaying raw values and instead support comparisons between different performance indicators.  </p>
<h2>3. KPI Tracking</h2>
<p>One of the main parts of the dashboard is its KPI layer.</p>
<p>The dashboard tracks <strong>7+ YouTube performance KPIs</strong>, giving users a quick overview of how the channel or videos are performing.</p>
<p>The idea was to make important information visible immediately rather than requiring users to inspect individual records.</p>
<h2>4. 30-Day Performance Analysis</h2>
<p>I focused on <strong>30-day video performance</strong> to understand how videos were performing over a recent period.</p>
<p>This provides a more useful analytical window than looking at isolated values.</p>
<p>The dashboard can therefore be used to explore recent performance patterns and identify how videos are behaving within the selected period.  </p>
<h2>5. Percentile-Based Comparison</h2>
<p>One of the more interesting parts of the project was comparing individual videos with <strong>typical channel performance using percentile-based insights</strong>.</p>
<p>Instead of asking only:</p>
<blockquote>
<p>"How many views did this video receive?"</p>
</blockquote>
<p>the analysis can ask:</p>
<blockquote>
<p>"How does this video's performance compare with what is typical for the channel?"</p>
</blockquote>
<p>This adds context to the raw metric.</p>
<p>A number becomes much more meaningful when you can understand where it sits relative to other observations.</p>
<h2>Visualization with Plotly</h2>
<p>To make the dashboard interactive, I used <strong>Plotly</strong>.</p>
<p>Interactive visualizations make it easier to explore performance data and switch between different analytical perspectives.</p>
<p>Rather than presenting static charts, the dashboard was designed to make the data itself more explorable.  </p>
<h2>Building the Dashboard with Streamlit</h2>
<p>The final dashboard was built using <strong>Streamlit</strong>.</p>
<p>Streamlit made it possible to turn the Python-based analysis into an interactive application without building a traditional frontend from scratch.</p>
<p>The final workflow became:</p>
<pre><code class="language-plaintext">YouTube Performance Data
          ↓
     Data Processing
          ↓
    Feature Engineering
          ↓
     KPI Calculation
          ↓
  Performance Analysis
          ↓
 Interactive Plotly Charts
          ↓
    Streamlit Dashboard
</code></pre>
<p>Tech Stack</p>
<h3>Programming &amp; Analysis</h3>
<ul>
<li><p>Python</p>
</li>
<li><p>Pandas</p>
</li>
<li><p>NumPy</p>
</li>
</ul>
<h3>Visualization</h3>
<ul>
<li>Plotly</li>
</ul>
<h3>Dashboard</h3>
<ul>
<li>Streamlit</li>
</ul>
<h1>What I Personally Built</h1>
<p>This project gave me hands-on experience across multiple stages of the analytics workflow.</p>
<p>I worked on:</p>
<ul>
<li><p>Preparing and processing the dataset</p>
</li>
<li><p>Performing feature engineering</p>
</li>
<li><p>Analyzing 30-day video performance</p>
</li>
<li><p>Building KPI-based analysis</p>
</li>
<li><p>Implementing percentile-based comparisons</p>
</li>
<li><p>Creating interactive Plotly visualizations</p>
</li>
<li><p>Developing the Streamlit dashboard</p>
</li>
</ul>
<p>Rather than using the dashboard only as a visualization layer, I worked on the analytical workflow that feeds the dashboard.  </p>
<h2>Project Links</h2>
<p><strong>GitHub:</strong> <a href="https://github.com/mahii-17/yt-dashboard">https://github.com/mahii-17/yt-dashboard</a></p>
]]></content:encoded></item><item><title><![CDATA[Customer Shopping Behavior Analysis: From Raw Transactions to Customer Insights]]></title><description><![CDATA[Introduction
Data analysis is not just about creating charts.
Before a dashboard can answer useful business questions, the underlying data needs to be clean, consistent, and structured in a way that m]]></description><link>https://sonuhere.hashnode.dev/customer-shopping-behavior-analysis-from-raw-transactions-to-customer-insights</link><guid isPermaLink="true">https://sonuhere.hashnode.dev/customer-shopping-behavior-analysis-from-raw-transactions-to-customer-insights</guid><dc:creator><![CDATA[Sonu Gupta]]></dc:creator><pubDate>Sun, 30 Aug 2026 08:17:48 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a93dde20fe06c4affc21105/3b3a2717-5abe-495d-835f-73e39c74a5ae.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Introduction</h2>
<p>Data analysis is not just about creating charts.</p>
<p>Before a dashboard can answer useful business questions, the underlying data needs to be clean, consistent, and structured in a way that makes analysis possible.</p>
<p>For this project, I worked with a retail customer dataset containing 3,900+ transactions and 18 features covering customer demographics, purchasing behavior, and product attributes.</p>
<p>The goal was to transform this raw transactional data into a reliable analytical dataset and use it to understand customer purchasing behavior.</p>
<p>The project workflow involved:</p>
<p>Data Cleaning → Feature Engineering → SQL Analysis → Power BI Visualization</p>
<h2>The Problem</h2>
<p>Retail transaction data can contain inconsistencies that make analysis difficult.</p>
<p>Before analyzing the data, I needed to address issues such as:</p>
<p>Missing values Inconsistent column names Redundant fields Data that was not immediately suitable for analysis</p>
<p>Simply loading the raw dataset into a dashboard would not guarantee meaningful results.</p>
<p>So, the first objective was to improve the quality of the dataset before performing deeper analysis.</p>
<h2>Dataset</h2>
<p>The dataset contains 3,900+ retail transactions with 18 features.</p>
<p>The available information covers areas such as:</p>
<p>Customer demographics Product attributes Purchasing behavior Transaction-related information</p>
<p>This combination made the dataset suitable for exploring customer segmentation and purchasing patterns.</p>
<h2>My Approach</h2>
<p>I divided the project into several stages.</p>
<ol>
<li>Data Cleaning with Python</li>
</ol>
<p>I started by examining the dataset and preparing it for analysis using Python and Pandas.</p>
<p>The cleaning process included:</p>
<p>Handling missing values Standardizing column names Removing redundant fields Checking the structure and consistency of the data</p>
<p>This preprocessing step improved the data quality by approximately 35% according to my project analysis.</p>
<p>This step was important because poor-quality input data can lead to misleading analysis later.</p>
<ol>
<li>Feature Engineering</li>
</ol>
<p>After cleaning the dataset, I created additional features that could provide more useful analytical perspectives.</p>
<p>Two examples were:</p>
<p>Age Groups</p>
<p>Instead of treating every customer age as an isolated value, customers could be grouped into meaningful age categories.</p>
<p>Purchase Frequency</p>
<p>I also engineered purchase-frequency information to help understand customer purchasing behavior and support better segmentation.</p>
<p>These transformations made the dataset more useful for analysis beyond the original raw columns.</p>
<ol>
<li>SQL &amp; MySQL Analysis</li>
</ol>
<p>Once the dataset was prepared, I used SQL and MySQL to work with the structured data and perform analytical queries.</p>
<p>This part of the project helped me practice:</p>
<p>Querying relational data Filtering records Aggregating information Grouping customers and transactions Extracting patterns from transactional data</p>
<p>Using SQL alongside Python also helped me understand how the same analytical problem can be approached from different parts of a data workflow.</p>
<ol>
<li>Power BI Dashboard</li>
</ol>
<p>After preparing and analyzing the data, I used Power BI to turn the results into a visual format.</p>
<p>The purpose of the dashboard was to make customer and purchasing information easier to explore and interpret.</p>
<p>Instead of looking at thousands of individual transactions, the dashboard provides a higher-level view of the dataset and helps communicate the results more clearly.  </p>
<p>Tech Stack Programming &amp; Analysis Python Pandas Database MySQL SQL Visualization Power BI Workflow Raw Retail Data ↓ Data Cleaning ↓ Data Transformation ↓ Feature Engineering ↓ SQL / MySQL Analysis ↓ Power BI Dashboard What I Learned</p>
<p>This project taught me that a good analytics project is not simply about making visualizations.</p>
<p>The quality of the final result depends heavily on everything that happens before the visualization stage.</p>
<p>Through this project, I strengthened my understanding of:</p>
<p>Data Cleaning</p>
<p>Real datasets often require significant preparation before they become analysis-ready.</p>
<p>Feature Engineering</p>
<p>Creating useful derived features can reveal perspectives that are not immediately available in the original dataset.</p>
<p>SQL</p>
<p>SQL provides a powerful way to query and analyze structured relational data.</p>
<p>Data Visualization</p>
<p>A dashboard should communicate information clearly rather than simply display as many charts as possible.</p>
<h2>End-to-End Analytics Workflow</h2>
<p>Working across Python, SQL/MySQL, and Power BI helped me understand how different tools can fit together in a single analytics workflow.</p>
<p>Key Takeaway</p>
<p>The biggest lesson from this project was simple:</p>
<p>Good analysis starts with good data.</p>
<p>Cleaning and transforming the dataset was just as important as the final dashboard.</p>
<p>By taking the project from raw transactions to a structured analytical dataset and finally to a visualization layer, I was able to practice a complete data analytics workflow rather than working on isolated tasks.</p>
<h3>Project Links</h3>
<p>GitHub: <a href="https://github.com/mahii-17/customer-behavior">Customer Shopping behavior</a></p>
<p>Portfolio: <a href="https://sonuhere.vercel.app/">Click Here</a></p>
]]></content:encoded></item></channel></rss>