<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:cc="http://cyber.law.harvard.edu/rss/creativeCommonsRssModule.html">
    <channel>
        <title><![CDATA[Stories by Sergey Abakumoff on Medium]]></title>
        <description><![CDATA[Stories by Sergey Abakumoff on Medium]]></description>
        <link>https://medium.com/@sAbakumoff?source=rss-13832a65f7f1------2</link>
        <image>
            <url>https://cdn-images-1.medium.com/fit/c/150/150/0*gUuxwJr_Oqw_Wa3A.jpg</url>
            <title>Stories by Sergey Abakumoff on Medium</title>
            <link>https://medium.com/@sAbakumoff?source=rss-13832a65f7f1------2</link>
        </image>
        <generator>Medium</generator>
        <lastBuildDate>Fri, 09 Oct 2026 07:15:09 GMT</lastBuildDate>
        <atom:link href="https://medium.com/@sAbakumoff/feed" rel="self" type="application/rss+xml"/>
        <webMaster><![CDATA[yourfriends@medium.com]]></webMaster>
        <atom:link href="http://medium.superfeedr.com" rel="hub"/>
        <item>
            <title><![CDATA[One More Analysis of GitHub and StackOverflow Data with Google BigQuery]]></title>
            <link>https://medium.com/hackernoon/catalog-of-references-to-stackoverflow-questions-found-in-github-sources-134415b97ecb?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/134415b97ecb</guid>
            <category><![CDATA[bigquery]]></category>
            <category><![CDATA[stackoverflow]]></category>
            <category><![CDATA[google-cloud-platform]]></category>
            <category><![CDATA[open-data]]></category>
            <category><![CDATA[github]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Mon, 30 Jan 2017 06:51:20 GMT</pubDate>
            <atom:updated>2019-05-17T10:06:46.630Z</atom:updated>
            <content:encoded><![CDATA[<p><strong>TL;DR </strong>I built the web-site where you can explore the Stack Overflow questions referenced in the source code in <a href="https://hackernoon.com/tagged/github">Github</a>. Check it out in <a href="http://sociting.biz">http://sociting.biz</a></p><p><strong>Motivation</strong><br>I am a big fan of <a href="https://hackernoon.com/tagged/google">Google</a> Cloud Platform, especially I love its data warehouse implementation called BigQuery. In summer of 2016 Github and Google made the open-source data available for everyone in BigQuery, here are the mind boggling numbers:</p><blockquote>This 3TB+ dataset comprises the largest released source of GitHub activity to date. It contains a full snapshot of the content of more than 2.8 million open source GitHub repositories including more than 145 million unique commits, over 2 billion different file paths, and the contents of the latest revision for 163 million files, all of which are searchable with regular expressions.</blockquote><p>I have since then never tired of exploring these data, revealing interesting patterns or extreme samples and publishing <a href="https://medium.com/@sAbakumoff">articles</a> about my findings.<br>In December of 2016 Google team has woken up my “researcher within” again — they’ve added Stack Overflow’s history of questions and answers to the collection of public datasets on BigQuery. In practice that means that the most popular programming chat in the world is now can be analyzed with the power of Google Cloud Platform, for example one can run the <a href="https://medium.freecodecamp.com/always-end-your-questions-with-a-stack-overflow-bigquery-and-other-stories-2470ebcda7f#.bhdeolgrt">sentiment analysis</a> on the Stack Overflow data and find out that Python developers post the lowest percent of negative comments overall! What excites me the most though is the ability to join the Stack Overflow data with other <a href="https://cloud.google.com/bigquery/public-data/">publicly available data sets</a>. For example one can try to find out whether the weather can affect the probability of a Stack Overflow question to be answered by using the data from <a href="https://cloud.google.com/bigquery/public-data/noaa-gsod">NOAA dataset</a><a href="https://cloud.google.com/bigquery/public-data/noaa-gsod)(I">(I</a> am actually going to conduct this research soon).<br>In the <a href="https://cloud.google.com/blog/big-data/2016/12/google-bigquery-public-datasets-now-include-stack-overflow-q-a">introduction</a> to the Stack Overflow data availability <a href="https://medium.com/@hoffa">Felipe Hoffa</a> provided the sample of joining Github and Stack Overflow data to find out which are the most referenced Stack Overflow questions in the GitHub code — specifically, Javascript. It gripped my attention because I noticed a couple of limitations: <br>* The query searches only for stackoverflow.com/questions/([0–9]+)/ pattern in the source code. However, there are alternative forms of referencing questions : it could a short form stackoverflow.com/q/([0–9]+)/ and it could be the direct reference to one of the answers, like stackoverflow.com/answers/([0–9]+)/<br>* The query deals only with JavaScript sources, but there are plenty of other programming languages.</p><p>So, I set out to build the catalog of the stack overflow questions referenced in the GitHub sources for popular programming languages.</p><p><strong>Getting the data</strong><br><em>Step 1</em> Finding lines of code in Github Sources that have references to StackOverflow questions or answers. <a href="https://bigquery.cloud.google.com/dataset/fh-bigquery:github_extracts">contents_top_repos_top_langs</a> table that keeps contents for the top languages from the top repos was kindly provided by Felipe Hoffa</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/b36290ba089bbb697e6cb0dc95cb574d/href">https://medium.com/media/b36290ba089bbb697e6cb0dc95cb574d/href</a></iframe><p>The result has been saved in the new table called <strong>so_ref_top_repos_top_langs</strong> which contains the rows like</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/a57d7064192729450f4f3fb1ac2c1ff2/href">https://medium.com/media/a57d7064192729450f4f3fb1ac2c1ff2/href</a></iframe><p><em>Step 2</em> Join the result with the StackOverflow data. The query should handle both of the questions id’s and answers id’s extracted from the source code</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/67d9021b25a739fd9e93834df70ba714/href">https://medium.com/media/67d9021b25a739fd9e93834df70ba714/href</a></iframe><p>The result contains the rows the look like</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/276a2c00986a4aa1e785ebd688301ad4/href">https://medium.com/media/276a2c00986a4aa1e785ebd688301ad4/href</a></iframe><p>There were roughly 31K records like this one, the next question was on how to visualize them.</p><p><strong>Building the web-site</strong><br>First of all I moved the resulting data to the SQLite database by creating separate table for each programming language. Then I built the [web-site](<a href="http://sociting.biz">http://sociting.biz</a>) that allows to navigate through the data by switching between the languages and jump to the Github source code to check how the information from the questions/answers was applied in the specific scenarios. I also caught this opportunity to play with ASP.NET Core and implement the web-site on my Macbook Pro, without Windows being involved. The resulting application uses the cross platform ASP MVC Web API on the back end and react+redux on the front end. The source code is fully available in the <a href="https://github.com/sAbakumoff/SoCiting2">Github repo</a>.</p><blockquote><a href="http://bit.ly/Hackernoon">Hacker Noon</a> is how hackers start their afternoons. We’re a part of the <a href="http://bit.ly/atAMIatAMI">@AMI</a>family. We are now <a href="http://bit.ly/hackernoonsubmission">accepting submissions</a> and happy to <a href="mailto:partners@amipublications.com">discuss advertising &amp;sponsorship</a> opportunities.</blockquote><blockquote>To learn more, <a href="https://goo.gl/4ofytp">read our about page</a>, <a href="http://bit.ly/HackernoonFB">like/message us on Facebook</a>, or simply, <a href="https://goo.gl/k7XYbx">tweet/DM @HackerNoon.</a></blockquote><blockquote>If you enjoyed this story, we recommend reading our <a href="http://bit.ly/hackernoonlatestt">latest tech stories</a> and <a href="https://hackernoon.com/trending">trending tech stories</a>. Until next time, don’t take the realities of the world for granted!</blockquote><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fupscri.be%2Fdde502%3Fas_embed%3Dtrue&amp;dntp=1&amp;url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fupscri.be%2Fhackernoon%2F&amp;key=a19fcc184b9711e1b4764040d3dc5c07&amp;type=text%2Fhtml&amp;schema=upscri" width="800" height="400" frameborder="0" scrolling="no"><a href="https://medium.com/media/3c851dac986ab6dbb2d1aaa91205a8eb/href">https://medium.com/media/3c851dac986ab6dbb2d1aaa91205a8eb/href</a></iframe><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=134415b97ecb" width="1" height="1" alt=""><hr><p><a href="https://medium.com/hackernoon/catalog-of-references-to-stackoverflow-questions-found-in-github-sources-134415b97ecb">One More Analysis of GitHub and StackOverflow Data with Google BigQuery</a> was originally published in <a href="https://medium.com/hackernoon">HackerNoon.com</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[‘Member HTML Login Forms?]]></title>
            <link>https://medium.com/@sAbakumoff/member-html-login-forms-78c3ad3ba4e4?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/78c3ad3ba4e4</guid>
            <category><![CDATA[react]]></category>
            <category><![CDATA[simplicity]]></category>
            <category><![CDATA[forms]]></category>
            <category><![CDATA[redux]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Thu, 01 Dec 2016 10:18:22 GMT</pubDate>
            <atom:updated>2016-12-01T10:20:03.090Z</atom:updated>
            <content:encoded><![CDATA[<p>‘Member circa 2010 when we used a pretty simple approach to implement the Login function in our web-applications? It was so awesome!<a href="http://www.w3schools.com/html/html_forms.asp"> HTML Form elements</a> gave us everything we need: attributes to specify the server-side handler, input controls for username, password and “remember me” checkbox, submit button. <a href="https://developer.mozilla.org/en-US/docs/Web/Guide/HTML/Forms/Data_form_validation">HTML5 allowed</a> us to ensure that the fields are not empty without relying on JavaScript code. The implementation of a server-side handler of a form submitting depended on the underlying technology but it also was very simple — if a username or password were incorrect, a browser was instructed to paint the login form again along with the error message, otherwise it was redirected to one of the main pages of a web-app. The beauty was in that the client-side was fundamental — we simply didn’t need anything else. Indeed the frameworks like ASP.NET provided helpers à la <a href="https://msdn.microsoft.com/en-us/library/system.web.ui.webcontrols.login.aspx">Login control</a> but in the end it boils down to the same HTML forms.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/688/1*wv2QKvaAWtNOkxK41xoQig.png" /><figcaption>Sample of code from w3schools</figcaption></figure><p>If a prehistoric developer had been frozen in 2010 and then thawed out in 2016, he would discover the whole new world. Imagine that he met a modern react.js ninja who tried to explain the way we are using to implement login forms in 2016.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/459/1*31RYR6sGQqf6pfFA3ReGvg.png" /></figure><p><strong>React components and containers. </strong>We now wrap login forms in the React Components in order to integrate Login function into React apps. The sample below was copied from <a href="https://github.com/suranartnc/react-redux-universal-starter-kit/blob/master/src/shared/containers/LoginPage/LoginForm/index.js">React Redux Universal Starter Kit</a> project.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/65ed2411074cef67125ec86928056498/href">https://medium.com/media/65ed2411074cef67125ec86928056498/href</a></iframe><p>And it’s not enough! Some guy <a href="https://medium.com/@dan_abramov/smart-and-dumb-components-7ca2f9a7c7d0#.d48f3nplc">convinced us</a> to keep apart presentational and container components, so we bring this separation even for such a simple function as Login. It’s funny that there is no single recipe but a wide range of solutions for doing that, here are few examples:</p><ul><li>Using <a href="http://redux-form.com/6.2.0/docs/GettingStarted.md/">redux-form</a> which is the decorator that wraps a form in a Higher Order Component (whatever it means), like it’s done in the <a href="https://github.com/suranartnc/react-redux-universal-starter-kit/blob/master/src/shared/containers/LoginPage/LoginForm/index.js">full code</a> for the Login Form shown above.</li><li>Introducing a container for Login form, like the one in <a href="https://github.com/relax/relax/tree/master/lib/shared/screens/auth/screens/login">relax project</a>:</li></ul><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/21502013cd3407f7018757769ed43add/href">https://medium.com/media/21502013cd3407f7018757769ed43add/href</a></iframe><ul><li>Adding the “handler” of login form, then wrapping it with the help of react-redux, just like it’s done in <a href="https://github.com/lancetw/react-isomorphic-bundle/blob/master/src/shared/components/TWBLoginHandler.js">React Redux Universal (isomorphic) bundle project</a>:</li></ul><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/bd72e2d3bded9d73a10b6b1a5e238ca5/href">https://medium.com/media/bd72e2d3bded9d73a10b6b1a5e238ca5/href</a></iframe><p>I guess that at this point the facial expression of the prehistoric developer would be close to what is shown in the Toy Story picture, but there are more things to explain.</p><p><strong>Single Page Apps and Reactive State. </strong>We don’t let browser send the login form by using its native capabilities. Not anymore. We are obsessed with the idea of hosting the whole application on a single page and simultaneously pretending that it contains multiple pages by having total control over browser’s history which requires manual submitting of a login form’s data. Therefore we intercept the <a href="https://developer.mozilla.org/en-US/docs/Web/Events/submit">submit event</a> of a form and pass the entered username and password further to the login request, but in order to obtain these values we don’t refer to the corresponding DOM elements(as we would normally do with jQuery). No sir. It is indeed possible, but the <a href="https://facebook.github.io/react/docs/refs-and-the-dom.html">React documentation</a> calls it “escape hatch” and suggest not to overuse it. Instead we used so-called “<a href="https://facebook.github.io/react/docs/forms.html">controlled</a> text inputs” that on every keystroke raise the change event whose handler updates the state object which is used to re-paint the input box. The snippet of code below copied from <a href="https://github.com/gelatinous-toboggan/gelatinous-toboggan/blob/master/app/src/components/login.js">gelatinous-toboggan</a> project:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/b254bbda026f14c5f7111b11174765c5/href">https://medium.com/media/b254bbda026f14c5f7111b11174765c5/href</a></iframe><p>It seems like we overwrite the native browser’s implementation of the input element‘s behavior time and again(One more <a href="https://github.com/relax/relax/blob/master/lib/shared/screens/auth/screens/login/components/login.jsx">example</a>). Though some delinquents <a href="https://github.com/MRN-Code/coinstac/blob/master/packages/coinstac-ui/app/render/components/form-login.js">bypass this practice</a>:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/bf3095738a8b5f511bc0cc825f371f11/href">https://medium.com/media/bf3095738a8b5f511bc0cc825f371f11/href</a></iframe><p>Prehistoric developer: do you have the time machine in 2016 so that I can escape from this madness?</p><p>Don’t get me wrong, I am not trying to discredit react.js. It’s great for solving the problems that it was invented to solve effectively. The issue is react and redux concepts — presentational and container components, controlled forms, state management are applied for each and every piece of functionality, even for very simple one like Login form, which only contributes to the complexity of the code and its size and performance, but does not solve the problem more effectively. The art of keeping it simple which is the paramount quality for a professional developer gradually vanishes and I will be happy if this article or <a href="https://medium.com/@sAbakumoff/es6-is-great-until-its-not-f398339d0af6#.wpwms1zcs">previous one</a> makes one developer to consider alternative approach to building the complex structures with no apparent reasons for his next project. Thanks for reading!</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=78c3ad3ba4e4" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[ES6 is great, but use it cautiously]]></title>
            <link>https://medium.com/@sAbakumoff/es6-is-great-until-its-not-f398339d0af6?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/f398339d0af6</guid>
            <category><![CDATA[optimization]]></category>
            <category><![CDATA[javascript]]></category>
            <category><![CDATA[performance]]></category>
            <category><![CDATA[es6]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Wed, 16 Nov 2016 15:46:37 GMT</pubDate>
            <atom:updated>2016-11-17T07:13:28.611Z</atom:updated>
            <content:encoded><![CDATA[<p>The other day I came across a hilarious tweet</p><h3>Nick Williams on Twitter</h3><p>How I imagine this code came to be: &quot;Let&#39;s use fancy ES6 destructuring&quot; &quot;Seems hard to read&quot; &quot;Fear not! I&#39;ll leave an explanatory comment</p><p>It’s funny indeed, but at the same time it’s the typical showcase of Cargo Cult Programming — ritual inclusion of code or program structures that serve no real purpose. It might seem to be an edge case, a funky thing that couldn’t be really met in the real-world, but let’s review a piece of code from the highly rated — more than 8000 Github stars — project called <a href="https://github.com/Automattic/wp-calypso">Calypso</a>.</p><blockquote>Calypso is the new WordPress.com front-end — a beautiful redesign of the WordPress dashboard using a single-page web application, powered by the WordPress.com REST API. Calypso is built for reading, writing, and managing all of your WordPress sites in one place.</blockquote><p>It’s clearly on the cutting edge — react, redux, webpack, babel and <a href="https://github.com/Automattic/wp-calypso/blob/master/package.json">other 150 npm packages</a> it depends on should allow contributors to write a modern, clean, easily readable, highly maintainable code, right? Let’s take a look. But first things first, here is a crash course on redux for those who are not familiar with it — the <a href="http://redux.js.org/docs/basics/Reducers.html">reducer</a> is a <a href="https://en.wikipedia.org/wiki/Pure_function">pure function</a> that takes two objects, performs some calculation and returns the new object, that’s basically it. A primer on reducer is adding a new item in the TODO list :</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/b5297b9ce4d3ae2fa85a439a33cd8e6f/href">https://medium.com/media/b5297b9ce4d3ae2fa85a439a33cd8e6f/href</a></iframe><p>Multiple reducers can be composed within the single object where the keys are reducers names and the values are reducers functions. Here is the slightly modified code for one of the <strong>many</strong> Calypso reducers, extracted from <a href="https://github.com/Automattic/wp-calypso/blob/master/client/state/sharing/publicize/reducer.js">here</a>.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/1edd7505f2f8c1f3b47491d0e81da818/href">https://medium.com/media/1edd7505f2f8c1f3b47491d0e81da818/href</a></iframe><p>How long did it take for you to figure out what exactly is going on here? If you grasped it instantly then congratulations, you are true javascript ninja, because boy, oh boy, this code really gets the most out of ES6 and even ES7!</p><ul><li><a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Statements/const">const declarations</a></li><li><a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Operators/Object_initializer#Computed_property_names">computed property names</a></li><li><a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Operators/Destructuring_assignment">Pulling fields from objects passed as function parameter</a></li><li><a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Functions/Arrow_functions">Arrow functions</a></li><li><a href="https://github.com/sebmarkbage/ecmascript-rest-spread/blob/master/Spread.md">Spread Object properties</a> — it’s not even in the specification yet, but we can use it today, how cool is that!</li><li><a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Operators/Object_initializer#Property_definitions">Shorthand Property Declarations</a></li></ul><p>But to me — and I guess that I’m not alone — this piece-of-art seems too fancy, too difficult to understand, and if I were to implement the same functionality, I would do it slightly differently. This reducer updates the state of a given post publication on a given web-site, here we go:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/1687acb6e6fc4101f601eb5a5f4670f4/href">https://medium.com/media/1687acb6e6fc4101f601eb5a5f4670f4/href</a></iframe><p>And if you think that it’s just a matter of taste, let me show that there are more serious issues with the fancy version of code. First of all, both of the original version and the modification need to be “compiled” to javascript that is supported in old browsers(IE9 or Android Stock), <a href="https://babeljs.io/repl/">the Babel REPL</a> generates the output of the following sizes(click “full” word to review the compiled code).</p><ul><li>Original — <a href="https://gist.github.com/sAbakumoff/2b0ed1699ab6020e1b49faacca483f3f">full</a> : 2012 bytes, minified : 1524 bytes</li><li>Modified — <a href="https://gist.github.com/sAbakumoff/14118c1fe71d2810c8615ac01d12a8d3">full</a>: 1390 bytes, minified : 1019 bytes</li></ul><p>The difference might look minor, but in a large project, like Calypso, hundreds if not thousands lines of code written in the name of Cargo Cult might add up to the significant increase in JS scripts size which is one of the key points to avoid in order to optimize the <a href="https://developers.google.com/web/fundamentals/performance/critical-rendering-path/">critical rendering path</a>.</p><p>Moreover, the modified version of a reducers works FASTER. I put together <a href="https://jsperf.com/fancy-vs-humble">JSPerf test</a> to demonstrate that, here is the screenshot of test run in Chrome:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DFQhVzYg0Vji15dHbbrxRw.png" /></figure><p>Note that modified version of the code also leverages ES6 features, but it does not try to use each and every buzz-thing in the name of fanciness. In the end the code is more readable, faster and smaller. And I can optimize it even further! Let me finish this post with one of my favorite quotes about programming.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/638/1*E6x7IEaGSA3Tay1pHRjX2Q.png" /></figure><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=f398339d0af6" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[React Entourage]]></title>
            <link>https://medium.com/@sAbakumoff/react-entourage-6d51e7df9944?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/6d51e7df9944</guid>
            <category><![CDATA[github]]></category>
            <category><![CDATA[javascript]]></category>
            <category><![CDATA[bigquery]]></category>
            <category><![CDATA[stats]]></category>
            <category><![CDATA[react]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Sat, 29 Oct 2016 07:30:55 GMT</pubDate>
            <atom:updated>2016-10-30T15:54:45.745Z</atom:updated>
            <content:encoded><![CDATA[<p>So, you’ve decided to use react.js in your new awesome project and now you’ve got a new problem — which state management approach to choose? Is it going to be Redux, Fluxible or Alt? As you work on a project, <a href="https://github.com/enaqx/awesome-react">a ton</a> of 3rd party tools, libraries and components are right there to help. <a href="https://cloud.google.com/bigquery/public-data/github">Public Github data</a> allows to look at the summary stats on this flourishing field of modern react-based applications.</p><h3>Getting and Cleaning Data</h3><p>The most common approach to build a web app nowadays is bundling a bunch of npm packages — react, redux, react-redux, etc. to be used in a browser environment with a help of tools like Gulp or Grunt. These packages are often listed in the “package.json” file, so I run the following query to select public Github repositories that depend on React. It is optimized to include package.json files that locate in the root folder of a repository only.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/c93d159b0b0c1b90eb43347d5c5d74ca/href">https://medium.com/media/c93d159b0b0c1b90eb43347d5c5d74ca/href</a></iframe><p>The resulting data table contains all the run-time(isDev=0) and design-time(isDev=1) dependencies for each repository that relies on “react”:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/519/1*9R0i8P3P_On3jcY1Pc_NrA.png" /></figure><p>But these are raw data — some repositories could have been removed, others might contain meaningless code and so on. To filter it out, I run the <a href="https://github.com/sAbakumoff/react-entourage/blob/master/gh-data.js">script</a> that obtains the number of stars, forks, watchers and size of each distinct repository from the list by using the public Github API and saves the result into a <a href="https://github.com/sAbakumoff/react-entourage/blob/master/repos.csv">CSV file</a>:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/758/1*QHSpz-WnxFWYLgpYOqQXgA.png" /></figure><p>There are two types of projects in this list :</p><ul><li>Auxiliary tools, libraries, boilerplates and frameworks, like <a href="https://github.com/jumpsuit/jumpsuit">zab/jumpsuit</a>:</li></ul><blockquote>A powerful and efficient Javascript framework that helps you build great apps. It is the fastest way to write scalable React/Redux with the least overhead.</blockquote><figure><img alt="" src="https://cdn-images-1.medium.com/max/316/1*pAfpBxlP5yUG5jzdAt3rFg.gif" /></figure><ul><li>Applications that rely on react and probably on these tools, like <a href="https://github.com/Bobeta/reactive-weather">Reactive Weather</a>(pretty cool stuff).</li></ul><p>The target of this research is the second type — applications. They could be separated by assuming that other projects do not depend on them:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/b9323d7b6e17639f61b2db17c173b482/href">https://medium.com/media/b9323d7b6e17639f61b2db17c173b482/href</a></iframe><h3>Deriving the stats</h3><p>The other day I came across a tweet that made fun of JS apps</p><h3>Joe Groff on Twitter</h3><p>JavaScript 4K demo competition! Build an amazing app with fewer than 4096 dependencies</p><p>here is the distribution of number of dependencies among react-based ones:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/c7ff717484fb1567d705b7bfa830ed97/href">https://medium.com/media/c7ff717484fb1567d705b7bfa830ed97/href</a></iframe><p>75% of projects have no more than 20 dependencies and 24 devDependencies, that’s not bad at all, though the nested dependencies may add up to 4096. Here are several outliers that could achieve this number.</p><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2Fembed%2FshrZuA6s3ZUfT8VQq&amp;url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2FshrZuA6s3ZUfT8VQq&amp;image=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fstatic.airtable.com%2Fimages%2Foembed%2Fairtable.png&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;type=text%2Fhtml&amp;schema=airtable" width="800" height="533" frameborder="0" scrolling="no"><a href="https://medium.com/media/55ff64865bde0cf890b3c957d625e922/href">https://medium.com/media/55ff64865bde0cf890b3c957d625e922/href</a></iframe><p>Top 100 run-time dependencies. It seems that redux really holds the dominant position there. It’s also interesting that jquery is still in the game and even outperforms, in terms of number of usages, material-ui package. Good job jQuery, I will never be disappointed with you.</p><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2Fembed%2FshrfmD9l5rAW4Z03N&amp;url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2FshrfmD9l5rAW4Z03N&amp;image=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fstatic.airtable.com%2Fimages%2Foembed%2Fairtable.png&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;type=text%2Fhtml&amp;schema=airtable" width="800" height="533" frameborder="0" scrolling="no"><a href="https://medium.com/media/3b0039592f66839f8321c17a95e5bfb6/href">https://medium.com/media/3b0039592f66839f8321c17a95e5bfb6/href</a></iframe><p>The same for design-time dependencies:</p><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2Fembed%2FshrifN8EOqEYyjULo&amp;url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2FshrifN8EOqEYyjULo&amp;image=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fstatic.airtable.com%2Fimages%2Foembed%2Fairtable.png&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;type=text%2Fhtml&amp;schema=airtable" width="800" height="533" frameborder="0" scrolling="no"><a href="https://medium.com/media/5d0d0e4f1311b7d7b4d5c6757fa9072e/href">https://medium.com/media/5d0d0e4f1311b7d7b4d5c6757fa9072e/href</a></iframe><p>Finally, I saw <a href="https://news.ycombinator.com/item?id=12802121">a request</a> for good samples of applications written in React, so I compiled <a href="https://gist.github.com/sAbakumoff/7b8510adcb16bded189d747e34f5e114">the list</a> of top 1000 projects arranged by the number of stars. I sure that it’s a good material for learning the ways React is used in the real-world.</p><p>That’s all dear reader, hope that the article was interesting to read. Stay tuned!</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=6d51e7df9944" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[A Deeper Dive Into JavaScript Arrays]]></title>
            <link>https://medium.com/@sAbakumoff/a-deeper-dive-into-javascript-arrays-390f324f7482?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/390f324f7482</guid>
            <category><![CDATA[javascript]]></category>
            <category><![CDATA[ecmascript]]></category>
            <category><![CDATA[tips-and-tricks]]></category>
            <category><![CDATA[arrays]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Sun, 09 Oct 2016 11:36:26 GMT</pubDate>
            <atom:updated>2016-10-09T11:36:26.683Z</atom:updated>
            <content:encoded><![CDATA[<p>The 5th edition of ECMA-262 standard aka ES5 introduced really useful methods of the built-in <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Array">Array object</a> : <em>forEach</em>, <em>every</em>, <em>some</em>, <em>filter</em>, <em>map</em>, <em>reduce</em>, <em>reduceRight</em>. The only problem is that they may not be present in all implementations of the standard, for example in MS IE8 implementation that is based on the 3rd edition of the ECMA-262. Fortunately we can work around this by inserting the polyfill for these methods in our scripts. It’s all well and good, but a polyfill’s code might look a little bit odd, for example here is the slightly modified <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Array/forEach#Polyfill">MDN implementation</a> of <em>forEach </em>method<em>:</em></p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/5cd0a2be2b42e15807a7c40cae2f8867/href">https://medium.com/media/5cd0a2be2b42e15807a7c40cae2f8867/href</a></iframe><p>The loop on lines 13–18 seems a little bit strange, doesn’t it? On the surface it may be unclear why the implementation isn’t as simple as something like this:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/9da5bb8cd67216f9eb8ee10d03559c71/href">https://medium.com/media/9da5bb8cd67216f9eb8ee10d03559c71/href</a></iframe><p>But it turns out that this “simple” polyfill may cause significant performance penalty because ECMAScript lets developers create “sparse” arrays that have arbitrary length, just like this:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/c6ac77a709a33f54aada154d8abb27b5/href">https://medium.com/media/c6ac77a709a33f54aada154d8abb27b5/href</a></iframe><p>The “simple” forEach polyfill called on a1, a2 or a3 array will invoke <em>callback</em> function <strong>1B+1</strong> times including <strong>1B redundant calls</strong> for the elements of the array which do not actually exist, for example a1[0]…a1[1000*1000*1000–1] are missing and return <em>undefined </em>value. Not only it could mislead the callback’s code if it handles undefined values, for example counts how many of them exist in a given array, but it also may significantly increase the running time of a code. So this is completely wrong way to implement the polyfill of <em>forEach </em>method, moreover, you better be vigilant about using traditional <strong><em>for(var i=0;i&lt;a.length;i++)</em></strong><em> </em>loop<em> </em>for traversing arrays that could come up to your script from an external, potentially <a href="https://blog.chibicode.com/you-can-submit-a-pull-request-to-inject-arbitrary-js-code-into-donald-trumps-site-here-s-how-782aa6a17a56#.k2wxe87lj">harmful</a> code. Fortunately the <a href="http://es5.github.io/#x15.4.4.18">ES5 standard</a> explicitly specifies how <em>forEach, map, filter, etc </em>solve this problem:</p><blockquote><strong>callbackfn is called only for elements of the array which actually exist; it is not called for missing elements of the array.</strong></blockquote><p>The MDN polyfill gracefully implements this requirement by leveraging the <a href="http://es5.github.io/#array-element">distinctive feature</a> of Array objects properties:</p><blockquote>Array objects give special treatment to a certain class of property names. A property name P (in the form of a String value) is an array index if and only if ToString(ToUint32(P)) is equal to P and ToUint32(P) is not equal to 232−1.</blockquote><p>In other words if an array has property with name “42” then the element with index 42 does exist in the array and vice versa. The properties are adjusted automatically when an array is being modified. Thus, <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Operators/in"><strong>in operator</strong></a> that is called in line 14 of the MDN polyfill checks whether an element with index <em>k</em> actually exists in an array. Here is the descriptive example of using this approach:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/6b8600a4ed6c4e0d5df4036784f83515/href">https://medium.com/media/6b8600a4ed6c4e0d5df4036784f83515/href</a></iframe><p>So far, so good, but here is another interesting observation : the polyfill does not check if it was called on the actual array, moreover the specification says that <em>forEach, map, filter, etc.</em> function:</p><blockquote>is <strong>intentionally generic</strong>; it does not require that its <strong>this</strong> value be an Array object. Therefore it can be transferred to other kinds of objects for use as a method.</blockquote><p>Technically, a method of one object can be “transferred” to another object by using <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Function/call">Function.prototype.call</a>. The most well-known application of this feature is calling Array methods on Array-like objects:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/68037bbd8e57228077226ba48a51b294/href">https://medium.com/media/68037bbd8e57228077226ba48a51b294/href</a></iframe><p><a href="https://developer.mozilla.org/en/docs/Web/JavaScript/Reference/Functions/arguments"><em>arguments</em> object</a> that corresponds to the parameters passed to a function is not an Array, but it has <em>length</em> property and indexed elements, hence <em>Array.prototype.filter</em> function can perfectly handle it. Note that in the code above <em>filterFn.call</em> returns the new Array object, so the subsequent <em>forEach</em> call does not require transferring.</p><p>Another example of Array-like object is <a href="https://developer.mozilla.org/en/docs/Web/API/HTMLCollection">HTMLCollection</a>:</p><iframe src="https://cdn.embedly.com/widgets/media.html?url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fjsfiddle.net%2Fa6gcsm8c%2F&amp;src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fjsfiddle.net%2Fa6gcsm8c%2Fembedded%2F&amp;type=text%2Fhtml&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;schema=jsfiddle" width="600" height="400" frameborder="0" scrolling="no"><a href="https://medium.com/media/b3444654739849dd10832903d443a966/href">https://medium.com/media/b3444654739849dd10832903d443a966/href</a></iframe><p>And one more trick: ECMAScript 5 introduced the way to treat the string as an array-like object, where individual characters correspond to a numerical index. You probably use this feature all the time, but take a look at how it can be leveraged to reverse a string, I guess that you probably didn’t see anything like that before:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/f4b8c75098c3e673f8b9009231cc27d6/href">https://medium.com/media/f4b8c75098c3e673f8b9009231cc27d6/href</a></iframe><p>That’s all, dear reader, for the first episode of “A Deeper Dive Into JavaScript” series. Hope you enjoyed it. Feel free to contact me if you have any questions.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=390f324f7482" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Real-time Analytics with Google Cloud]]></title>
            <link>https://medium.com/@sAbakumoff/real-time-analytics-of-presidential-debate-feedback-with-google-cloud-2dc55bacb1da?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/2dc55bacb1da</guid>
            <category><![CDATA[debate]]></category>
            <category><![CDATA[analytics]]></category>
            <category><![CDATA[google-cloud-platform]]></category>
            <category><![CDATA[big-data]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Tue, 27 Sep 2016 08:08:12 GMT</pubDate>
            <atom:updated>2016-09-29T02:16:28.847Z</atom:updated>
            <content:encoded><![CDATA[<p>The Internet era changed the way we obtain and process information<strong>. </strong>The huge volumes of data are free and ubiquitous and the cloud computing has put practically infinite computing power and storage and the sophisticated tools at everyone’s disposal, on a pay-as-you-go basis. This story explains how to leverage the cloud computing power and publicly available data to build a tool for real-time analytics of the presidential debate feedback posted to twitter. The aim of this article is to show how easy it is to implement pretty interesting, helpful(or malicious!) and even lucrative applications by using the Google Cloud Platform which is the quintessence of the Internet Age Marvel.</p><h4>Architecture</h4><p>Among other mind-blowing services such as <a href="https://cloud.google.com/bigquery/">BigQuery</a> or <a href="https://cloud.google.com/natural-language/">Natural Language API</a>, Google provides a set of tools for super-fast information exchange and data processing.</p><ul><li><a href="https://cloud.google.com/pubsub/">Pub/Sub</a> : real-time communication service that allows to send up to 10K messages/sec, ensures encryption, guarantees delivery and operates in the Google infrastructure.</li><li><a href="https://cloud.google.com/dataflow/">DataFlow</a> : real-time auto-scaling data processing service for building fast, reliable and secure Extract-Transform-Load pipelines that are integrated with with other Google services. For example a pipeline can receive the messages from Pub/Sub service, process them by using the Natural Language API and save the outcome in a BigQuery data set.</li></ul><p>Another player in the ensemble I put together is the <a href="https://dev.twitter.com/streaming/overview">Twitter Streaming API</a> that makes it possible to intercept the real-time stream of tweets that are sent from a particular location, by a specific user or contain certain phrases, for example “Hillary Clinton” or “Donald Trump”. At this point you probably figured out the whole picture as it’s quite simple, but just in case here is the sequence diagram of the architecture that performs the sentiment and syntax analysis of a stream of tweets and saves the results in a BigQuery data set. All of it, except of the Twitter back-end and client app(which is pretty simple) happens at real-time in the Google Cloud.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ng8ltmQ-zeXikLWpYboynw.png" /></figure><h4>Implementation</h4><p>I believe that the key point of this story is the incredible ease of the code that orchestrates the system shown in the picture above. The sources are available in the <a href="https://github.com/sAbakumoff/InFullGear">Github repository</a> and you can use it under the MIT license terms. I only wrote <strong>only 150 lines of code and copy-pasted 200 more lines. </strong>For example take a look at the node.js function that connects to a twitter stream of debate-related messages and sends the <a href="https://dev.twitter.com/overview/api/tweets">tweets objects</a> automatically serialized to JSON strings to the Pub/Sub service as soon as they are delivered. Note that the code is only interested in the tweets that are sent from the USA cities. This is done in order to make sure that the further data analysis could compare the data associated with different cities or states of the country that hosts the debate and elections. In addition it helps to prevent possible excess of the Natural Language API <a href="https://cloud.google.com/natural-language/limits">limits</a> that permit only 1000 requests per 10 seconds, but the Pub/Sub and DataFlow are able to handle 10K tweets per second!</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/9da07383325582e8034f256b97fd1bb8/href">https://medium.com/media/9da07383325582e8034f256b97fd1bb8/href</a></iframe><p>What is the “topic” argument passed in this function? In the Pub/Sub service a publisher sends messages to a <strong>topic</strong>. Subscribers components create a subscription to a topic to receive messages from it. Here is the code that creates the “debates_tweets” topic and passes it to startTracking function shown above:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/00cab394dc3541aa614412b762094650/href">https://medium.com/media/00cab394dc3541aa614412b762094650/href</a></iframe><p>The code that runs the DataFlow pipeline that subscribes to the “debate_tweets” topic, extracts a twitter text from the incoming message, processes it with the Natural Language API and saves the results in the BigQuery table is also quite simple, although it was quite an experience to write a code in Java that I didn’t touch for years.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/32d81ec0e287bc08307725b2a3bdac5d/href">https://medium.com/media/32d81ec0e287bc08307725b2a3bdac5d/href</a></iframe><p>The messages that the pipeline receives from the subscription are <a href="https://dev.twitter.com/overview/api/tweets">tweets objects</a> serialized to JSON strings. Note that the code does not try to fully parse them and then put the fields in the separate columns of a table row, but rather saves these strings as is in the single column. The rationale behind it is the data are saved in the BigQuery that allows to include the JavaScript functions in the SQL queries and parsing JSON with JS code is a cakewalk. Moreover, the code serializes the syntax analysis results to JSON strings and preserves them in the single column — I found it pretty convenient to keep the arrays of variable lengths that way. The sentiment analysis is represented with 2 float values that can be stored in two separates columns.</p><h3>Debate Night</h3><p>At 9pm EST September, 26 I started the pipeline and run the twitter stream client to harvest the tweets that were posted during the first presidential debate. 2 hours later the BigQuery table has been populated with 7K rows of data that include the information about the authors of the tweets, their location and the sentiment and syntax analysis results. I just needed a good tool for data exploration, analysis, and visualization and Google offers it as well! <a href="https://cloud.google.com/datalab/">Cloud Datalab</a> is the browser-based tool that can be used to interactively explore, transform, analyze, and visualize data using BigQuery and Python. The amount of data I’ve collected is not that big, but still the Datalab proved itself to be extremely helpful. Here is the sample of SQL+JS query that selects the most common nouns used in the tweets and its output.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/6f1e98614c077e7cbacdb8743e03a9ff/href">https://medium.com/media/6f1e98614c077e7cbacdb8743e03a9ff/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/823/1*ofkgshKQd11MYSLr2hffjQ.png" /></figure><p>You can find the full analysis notebook <a href="https://github.com/sAbakumoff/InFullGear/blob/master/data_analysis.ipynb">here</a>, Github nicely visualizes it. It’s just early, very basic steps though, the full analysis will take some time and it perhaps will be a subject of my next post, but the initial results are included in this story. First of all, here are some syntax analysis visualizations.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/568/1*vY5vSlo5J_BVZ6xK2yzm0w.png" /></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/539/1*5ETsvKUBpWIIdRVfEdikPA.png" /></figure><p>As for sentiment analysis, the Google API returns two float values for a given text : polarity and magnitude. Here is the explanation.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/869/1*9T_-ZFTGmRPeen0JfVdYIg.png" /></figure><p>Here are the plots of polarity and magnitude distribution for both candidates:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/780/1*OT0TdoT4JxbfgZHD7r3Pag.png" /></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/792/1*KJwPFCnuCbFPiry2VIt5wA.png" /></figure><p>It’s up to a reader to interpret these plots. What I can see here is a lot of negative feedback in the tweets about Mr. Trump.</p><h3>Imagine the unimaginable</h3><p>Let’s ponder for a minute on what could be possible to do by using the same technologies and even the very same code that was described in the article. This is actually the main purpose of this story, the debates stuff was used to get some attention ;)</p><ul><li>Startup idea : let’s say it’s Friday night in NewYork and a young hipster chooses a destination to go out. The app on his mobile phone shows the interactive map that highlights the places from where the most positive tweets are being sent. On the server side the app uses the real-time twitter analytics, just like the one that this story is about. In addition the app takes into account the analysis of photos that are shared via twitter. It is done by using the Google <a href="https://cloud.google.com/vision/">Vision API</a> that provides the insight from the images, <strong>including facial expressions of emotions</strong>!</li><li>Surveillance and Investigation. Twitter API allows to track the tweets from a specific user. By analyzing the info extracted from a stream of messages and photos that he or she posts on a daily basis it’s possible, for example, to predict(by using <a href="https://cloud.google.com/ml/">Google Cloud Machine Learning</a>) an attempt to commit a crime or suicide, or just watch a person’s life. If you are a fan of <a href="http://www.imdb.com/title/tt4158110/">“Mr. Robot”</a> TV show, you probably recall the Eliot’s manual hack of <a href="https://www.youtube.com/watch?v=OcePJlrAgSU">Fernando Vera</a>. I am pretty sure that it’s feasible to automate such a tedious task:</li></ul><blockquote>Fernando Vera, Shayla’s supplier. One of the worst human beings I’ve ever hacked. His password? “eatadick6969.” Aside from the massive amounts of money he spends on porn and webcams, he does all his drug transactions through emails, IMs, Twitter. The fact that the cops haven’t caught him yet is beyond me. If they had half a brain cell, they’d be able to crack his gang’s simplistic code, if it can even be called that. After only a <strong>couple of hours of timing his tweets with related news articles</strong>, I figured out that “biscuit” and “clickety” clearly referenced guns. “Food,” “sea shells” or “gas” for bullets. “rock to sleep early.” I haven’t made the direct connection to a hit yet, but the math of guns plus bullets usually adds up to one thing.</blockquote><ul><li>Real-time brand monitoring is a piece of cake that would take a day or two to implement.</li></ul><p>By the way, if you are a startup investor, private investigator or brand rep who are interested to collaborate, feel free to contact me. I will quit by boring job and we will watch the world together. Oppressive governments are not welcomed :p</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=2dc55bacb1da" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Angular vs React : text analysis of commit messages]]></title>
            <link>https://medium.com/@sAbakumoff/angular-vs-react-text-analysis-of-commit-messages-1cda199f3bdb?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/1cda199f3bdb</guid>
            <category><![CDATA[angularjs]]></category>
            <category><![CDATA[github]]></category>
            <category><![CDATA[react]]></category>
            <category><![CDATA[reactjs]]></category>
            <category><![CDATA[analysis]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Sun, 18 Sep 2016 06:06:02 GMT</pubDate>
            <atom:updated>2016-09-19T02:22:54.517Z</atom:updated>
            <content:encoded><![CDATA[<p>Every enthusiastic javascript blogger should write at least one article that compares the front-end frameworks, so do I! But no worries, this article isn’t Nth attempt to describe the pros and cons of the Angular and React, it rather glances at them from an unusual angle by applying the text mining methods to the commit messages that are collected from the version control system and remarking on the results. The research is fully reproducible, you can find the data and Rmd file <a href="https://github.com/sAbakumoff/frameworks-commits">here</a>. The details of the text analysis methods that were used in this article are described in the great book <a href="http://tidytextmining.com/">“Tidy Text Mining with R”</a> by Julia Silge and David Robinson.</p><h3>Getting data</h3><p>The first step in any data analysis is getting and cleaning data. In order to collect the React and Angular commit messages along with the relevant information, I used the <a href="https://cloud.google.com/bigquery/public-data/github">BigQuery public Github data</a>:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*pSsJShiccVn0LJRzL6wN_g.png" /></figure><p>The results have been saved in the table called “react_angular_commits” that has been exported to a <a href="https://github.com/sAbakumoff/frameworks-commits/blob/master/react_angular_commits.csv">CSV file</a> and downloaded for further processing that happened in R Studio. There are 7127 rows of the React messages and 8035 rows of the Angular messages.</p><h3>Warming up</h3><p>The first thing to compare is the length of the messages and drawing a <a href="https://en.wikipedia.org/wiki/Box_plot">box plot</a> seems to be very suitable for that, here is the corresponding R code and its output:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/82de9970f1d21b65f98d1dcb42d6633e/href">https://medium.com/media/82de9970f1d21b65f98d1dcb42d6633e/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/572/1*R8qvo-v4l6GXLU1miUS4Ng.png" /></figure><p>So, the median values are close(React : 81 characters, Angular: 74 characters), but it seems that the distribution of the Angular messages lengths is significantly skewed right compared with the React’s distribution. Simply speaking it means that the Angular contributors are more “voluble” when it comes down to write a commit message. The only reasonable explanation I came with is the difference in guidelines for contributors :</p><ul><li>Angular enforces pretty strict <a href="https://github.com/angular/angular.js/blob/master/CONTRIBUTING.md#commit">format of the commit messages</a>.</li><li>React <a href="https://github.com/facebook/react/blob/master/CONTRIBUTING.md">does not insist</a> on any particular format.</li></ul><p>But, I would expect the exactly opposite effect under these conditions — if everyone writes messages according to the pre-defined rules, then the length of these messages is not expected to vary a lot and vice versa. Apparently, it’s one of those cases where reality doesn’t meet expectations. Keep reading, more interesting stuff is coming :)</p><h3>Analyzing word and document frequency</h3><p>What are the most common words in the commits messages of Angular and React? To find it out these messages have to be split to individual words and cleaned off the <a href="https://en.wikipedia.org/wiki/Stop_words">stop words</a> with the help of <a href="https://cran.r-project.org/web/packages/tidytext">tydytext package</a>.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/f5610e4c35dd9c97d4537c80ab506cf8/href">https://medium.com/media/f5610e4c35dd9c97d4537c80ab506cf8/href</a></iframe><p>Here is the stats of Angular’s words frequency:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/7ef4190ac1db7e384b400aedc36ca092/href">https://medium.com/media/7ef4190ac1db7e384b400aedc36ca092/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/783/1*DuAWQ5F6witJfoZo400E9g.png" /></figure><p>Almost all of these words looks familiar to me except of <strong>“feat”</strong> and <strong>“chore”</strong>. The oxford dictionary provides the following definitions of these words:</p><blockquote><strong>feat</strong> — an achievement that requires great courage, skill, or strength</blockquote><blockquote><strong>chore</strong> — a routine task, especially a household one</blockquote><p>So, why are they so frequently used in the Angular’s commit messages? Let’s turn to the <a href="https://github.com/angular/angular.js/blob/master/CONTRIBUTING.md#commit">guidelines for contributors</a> for explanation:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/899/1*S7TJm-o3s38n_6Wtzq3Udw.png" /></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/897/1*W86EJcPLgf7G4sSv6_Cqxw.png" /></figure><p>Okay, so <strong>“feat”</strong> and <strong>“chore”</strong> are just types of commits! There are other types in the top 15 words : <strong>“test”</strong>, <strong>“fix”</strong> and <strong>“docs”</strong>. Hence the next idea — the header of a commit message also includes the scope of the changes, so let’s find out the most popular scopes along with the type of changes they are affected by. The following code might seem to be pretty convoluted, but it’s not — it extracts the type and scope from the Angular commit messages and saves the result to a new dataset that is joined with the most common scopes dataset in order to filter out the rest of scopes, finally the code plots the bar chart grouped by scope.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/bbcb008ac24d9caeba5e48c2ebb725e4/href">https://medium.com/media/bbcb008ac24d9caeba5e48c2ebb725e4/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DjUkB1VygbPwlzB2tjNABw.png" /></figure><p>That makes a lot of sense — for example, the most popular type of changes is “fix”, the documentation modifications affect the tutorial a lot and so on. One interesting observation : “$compile” scope demands a lot of attention in all the development areas. I don’t know a lot about $compile, but found a quite popular <a href="http://www.benlesh.com/2013/08/angular-compile-how-it-works-how-to-use.html">article</a> that says:</p><blockquote>View compilation in Angular is some of the most ingenious functional programming I’ve seen in JavaScript.</blockquote><p>That explains the amount of attention required! How cool is that? Simple text mining tools available to everyone allows to reveal what’s going on in the development of one of the most popular open source libraries!</p><p>Okay then, what about React’s most common words in the commit messages?</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/16621d5432618af0a90103c85a422999/href">https://medium.com/media/16621d5432618af0a90103c85a422999/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/709/1*pobfreYfYQnJm8FRKyWqow.png" /></figure><p>A lot of messages have “merge”, “pull” and “request” words. In fact 2274 React’s commit messages start with “Merge pull request #xxxx” string. The culprit behind it seems to be the <a href="https://help.github.com/categories/collaborating-on-projects-using-issues-and-pull-requests/">“Collaborating on projects using issues and pull requests”</a> workflow and in particular the recommended <a href="https://help.github.com/articles/merging-a-pull-request/">approach</a> to merge pull requests. The very same pattern pattern of commit messages can be found in other open-sources projects, for example take a look at <a href="https://github.com/rails/rails/commits/master">Rails</a> changes history. But it was not like that at all for the Angular commit messages! Why is it so? It seems that Angular maintainers use the alternative technique of integrating the contributors’ changes that boils down to applying the series of patches from the pull requests. You can find the detailed explanation in the excellent <a href="https://blog.spreedly.com/2014/06/24/merge-pull-request-considered-harmful/#.V9bDgj5945h">“Merge pull request Considered Harmful”</a> blog post. The Angular’s approach along with the rules for writing the commit messages turns the changes history to a useful product story that is easy to read and analyze. Unfortunately, React commits lack this beauty, but that’s not a reason to give up. There are other text mining methods that can tell a story of React, namely <a href="https://en.wikipedia.org/wiki/Tf%E2%80%93idf"><strong>term frequency–inverse document frequency</strong></a><strong> </strong>can be effectively used. Here is the code that selects the words that are frequently used in React’s messages, but rarely occurred in both of React and Angular messages.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/54381f58f06dc3c6431d87807ea49223/href">https://medium.com/media/54381f58f06dc3c6431d87807ea49223/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/695/1*TrKhPuPnrMaiBBOx2_-GZw.png" /></figure><p>This stats is much more informative than the previous than. First of all, “<a href="https://github.com/spicyj">spicyj</a>”, “<a href="https://github.com/zpao">zpao</a>”, “<a href="https://github.com/sebmarkbage">sebmarkbage</a>”, “<a href="https://github.com/chenglou">chenglou</a>”, “<a href="https://github.com/syranide">syranide</a>”, “<a href="https://github.com/jimfb">jimfb</a>”, “<a href="https://github.com/benjamn">benjamn</a>” refer to Github accounts of the most active <a href="https://github.com/facebook/react/graphs/contributors">contributors</a>. These names are used in the messages like</p><blockquote>Merge pull request #2343 from <strong>zpao</strong>/proptypes-deprecation Update PropTypes for ReactElement &amp; ReactNode</blockquote><p>so let’s ignore them and look at other highest tf-idf words. What about “korean” and “japanese” for example? It seems that React maintainers put a certain amount of efforts to translate the flagship <a href="https://facebook.github.io">documentation</a> to CJK. That’s was not observed for Angular and <a href="https://github.com/angular/angular.js/tree/master/docs/content/guide">its documentation</a> seems to have the English version only. Other common words indicate the areas that are frequently affected by the changes and every React developer should be familiar with them. Funny enough, they perfectly describe almost everything you need to implement and test a React-based application, here is the illustration:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1007/1*txzGmU51rb-AbNLnYTj2AA.png" /></figure><p>That’s all that I was able to extract from the React’s commit messages.Other methods like working with combinations of words did not help to obtain any other interesting information.</p><h3>Sentiment Analysis</h3><p>Finally let’s try to conduct the sentiment analysis of commit messages by using the dictionary that labels each word as “positive” or “negative”. Here is the code that plots the distribution of the most common positive and negative words in the angular’s messages:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/b8c4c61285d4f4d40ab819016177d88a/href">https://medium.com/media/b8c4c61285d4f4d40ab819016177d88a/href</a></iframe><figure><img alt="" src="https://cdn-images-1.medium.com/max/846/1*UknXSQ-9u3W8wITqsLsDoQ.png" /></figure><p>Indeed the words labeled as “negative” are not actually negative in context! We already know that “chore” is just a type of changes suggested in the guidelines for contributors. Or, for example, “error” is certainly not negative in the “fix ‘type mismatch’ error on IE8 after each request” message. The similar picture is observed for React:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/734/1*GfNjFBes9vKiKMiRTWngZg.png" /></figure><p>Practically it means that React and Angular contributors do not put a lot of emotion into the commit messages, for example there is only one <a href="https://github.com/facebook/react/commit/ec25297def5c6be933c064443250cfea02beab4c">f-bomb</a> among 15K commit messages. See my <a href="https://medium.com/@sAbakumoff/157-million-github-commits-48-thousand-f-bombs-a84cb9fab680#.2qd8ty3ud">previous post</a> for the examples of slightly different approach to formatting the commit messages ;)</p><h3>By way of conclusion</h3><p>In his great book “How Google Works” Eric Schmidt explains how astonishing things are in the Internet Century:</p><blockquote>Three powerful technology trends have converged to fundamentally shift the playing field in most industries. First, the Internet has made information free, copious, and ubiquitous — practically everything is online. Second, mobile devices and networks have made global reach and continuous connectivity widely available. And third, cloud computing has put practically infinite computing power and storage and a host of sophisticated tools and applications at everyone’s disposal, on an inexpensive, pay-as-you-go basis.</blockquote><p>This story is the example of using these powerful technologies : the data to analyze has been obtained from the publicly available source(Google Cloud) by using the free computational power(BigQuery) that tool only 15.6 sec to process 121Gb of data , the analysis has been done in the local machine, but I could’ve leverage the <a href="https://www.kaggle.com/kernels">kaggle kernels</a> platform to run the code in the cloud.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=1cda199f3bdb" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[156 Million Github commits, 48 Thousand F-bombs.]]></title>
            <link>https://medium.com/@sAbakumoff/157-million-github-commits-48-thousand-f-bombs-a84cb9fab680?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/a84cb9fab680</guid>
            <category><![CDATA[swearing]]></category>
            <category><![CDATA[bigquery]]></category>
            <category><![CDATA[big-data]]></category>
            <category><![CDATA[programming]]></category>
            <category><![CDATA[github]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Mon, 05 Sep 2016 12:59:17 GMT</pubDate>
            <atom:updated>2016-09-05T15:41:23.526Z</atom:updated>
            <content:encoded><![CDATA[<p>My Twitter friend <a href="https://medium.com/@hoffa">Felipe Hoffa</a> has recently posted the excellent story called <a href="https://medium.com/@hoffa/400-000-github-repositories-1-billion-files-14-terabytes-of-code-spaces-or-tabs-7cfe0b5dd7fd#.ml42l6ib2">400,000 GitHub repositories, 1 billion files, 14 terabytes of code: Spaces or Tabs?</a> which reveals the stats behind the never-ending-war between two styles of the code formatting. I was inspired by this article and got the idea for the new research : what’s the situation with swearing in the commit messages? Indeed it’s not new at all and there were multiple attempts to do the same :</p><ul><li><a href="http://geeksta.net/geeklog/exploring-expressions-emotions-github-commit-messages/">http://geeksta.net/geeklog/exploring-expressions-emotions-github-commit-messages/</a></li><li><a href="https://github.com/Dobiasd/programming-language-subreddits-and-their-choice-of-words">https://github.com/Dobiasd/programming-language-subreddits-and-their-choice-of-words</a></li></ul><p>But the recently <a href="https://cloud.google.com/bigquery/public-data/github">exposed BigQuery data</a> provides new opportunities to look at it. Also, as I am doing it mostly for fun, practicing my writing skills and contributing to the BigQuery community which I am a big fan of, the research is pretty simple : we will look at the frequency of dropping F-bombs in the commit messages and a couple of examples of such messages.</p><p>So, <a href="https://bigquery.cloud.google.com/table/bigquery-public-data:github_repos.commits">“commits” table</a> keeps the info about ~156 Million commits in the public repositories of Github, let’s select the ones that contain f-bombs in the subject or message:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*czmdWyUtZqV3qHeOYfIt2g.png" /><figcaption>Selecting commits that contain f-bombs in the subject and/or message</figcaption></figure><p>Reading through these results may cause a smile or two, for example(the comments in parentheses are mine ) :</p><ul><li>“Add some fucking class to #a Needs more Benedict Cumberbatch” (??)</li><li>“fix z-index, the motherfucker” (I know bro, z-index might be painful)</li><li>“no fuck this i’m done. i am so done.”(get some coffee)</li><li>“Rolled back license from “What the fuck you want” to the unlicense”(really?)</li><li>“WHAT THE HELL?! Who the fuck copied the CMS file to core folder and core file to CMS folder?!”(shit happens)</li><li>“Oh my fucking God I did a release with this debug code left in. HORROR.”(shit happens)</li><li>“switch to java 8… fuck backward compability”(fuck indeed:)</li><li>“Only that it needs to be the fucking opposite. Don’t drink and code…”(good idea!)</li><li>“No fucking idea what that will do”(classic)</li><li>“Fucking password is too fucking short”(fuck yeah)</li><li>“Hard, or Soft? So much to do. So decent salary. “You are unhappy? You do not want to stay here? Then there are bunches of people who are waiting to take the place of you.” You stay, or you fuck off. Choose one.”</li></ul><p>Okay, so how many f-bombs were dropped? To derive this numbers, I’ve saved the results shown above to the new <strong>“commits_with_f_bombs”</strong> table and run the next query, perhaps it’s not optimized, but it should be very easy to understand:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3yLPBUVbXm_K9mRoJVcB-Q.png" /></figure><p>So, only 48K of F-bombs, not bad. Indeed if I had expanded the queries to look for the <a href="https://en.wikipedia.org/wiki/Seven_dirty_words">7 dirty words</a> and their variations, the numbers would have been much larger :)</p><p>It would be also interesting to build the model of correlation between the programming language of the repository and the number of its commits containing f-bomb, but there are 2 problems here:</p><ul><li>The BigQuery data do not include information about the repository language.</li><li>The Github data actually has multiple languages per repository. For example <a href="https://github.com/ak72ti/Whynot">ak72ti/Whynot</a> which hosts 297 f-bomb commits:</li></ul><figure><img alt="" src="https://cdn-images-1.medium.com/max/999/1*-njkwpghlVh5-xe2mBJdkw.png" /><figcaption>Languages of repository</figcaption></figure><p>Which language caused the f-commits? What if the main(DM) is OK, but the contributors struggled with javascript part? Any ideas on this one is highly appreciated.</p><p>That’s it. Have a nice commit messages :)</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=a84cb9fab680" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[GitHub Cancer]]></title>
            <link>https://medium.com/@sAbakumoff/github-cancer-180db780d99d?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/180db780d99d</guid>
            <category><![CDATA[npm]]></category>
            <category><![CDATA[bigqu]]></category>
            <category><![CDATA[javascript]]></category>
            <category><![CDATA[open-source]]></category>
            <category><![CDATA[github]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Tue, 16 Aug 2016 14:52:21 GMT</pubDate>
            <atom:updated>2016-08-17T02:54:13.194Z</atom:updated>
            <content:encoded><![CDATA[<p>Earlier today I’ve sent the following tweet:</p><h3>I code therefore I&#39;m on Twitter</h3><p>&quot;isArray&quot; is new &quot;leftPad&quot; : ~25M downloads in the last month, ONE line of code : https://github.com/juliangruber/isarray/blob/master/index.js ... #javascript #wtf #npm #github</p><p>I’ve stumbled upon “isArray” thing during the exploration of the <a href="https://cloud.google.com/bigquery/public-data/github">public GitHub data</a> available in the Google BigQuery platform. This story exposes some new interesting findings I’ve discovered.</p><p>The previous stories(<a href="https://medium.com/@sAbakumoff/using-bigquery-github-data-to-rank-npm-repositories-ecf8947a1182#.v5f2euha0">1</a>, <a href="https://medium.com/@sAbakumoff/using-bigquery-github-data-to-find-out-open-source-software-development-trends-e288a2ca3e6b#.l2s6oskll">2</a>) analyzed the contents of package.json files which are descriptors used for Node.js modules dependency management. Recently, I noticed some odd numbers in the Github data:</p><ol><li>Number of files called “package.json” is over 8 Million.</li><li>Number of <em>contents</em> of the files called “package.json” is less than 1.5 Million.</li></ol><p>Huh? What is going on?</p><p>The answer is pretty simple — there is a huge number of <strong>duplicate</strong> package.json contents all over the open source code hosted in GitHub. Google BigData keeps the unique contents only, based on their hash I guess. The <a href="https://bigquery.cloud.google.com/table/bigquery-public-data:github_repos.contents">contents table</a> actually has “copies” column which indicates the number of duplicates, so let’s summarize the number of copies of package_json_content table that was <a href="https://medium.com/@sAbakumoff/using-bigquery-github-data-to-rank-npm-repositories-ecf8947a1182#.9yb7ixvn9">described</a> earlier:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*xxe-RK6UBaPb9LYy8ol5CQ.png" /><figcaption>Number of package.json files obtained from contents table</figcaption></figure><p>Cool, the total number of package.json files is now closer to the one from package_json_files table, they are not equal though, but let’s ignore it for now and move to the next query — what is the average number of the exact copies of a package.json file?</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*kxFuldTQhznbiiNM2-pohA.png" /><figcaption>Average number of duplicates</figcaption></figure><p>Let’s have a look at this phenomenon closer. I’ve composed and run the following query to extract the number of duplicates for each package.json file along with its NPM name, URL in github(one URL per each duplicate) and number of copies. The query selected the records with number of copies greater than 100 in the name of simplicity.</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/84bfb4c3d0fb8a06e0aa73a504749277/href">https://medium.com/media/84bfb4c3d0fb8a06e0aa73a504749277/href</a></iframe><p>The results have been saved in the new table called package_json_duplicates, then I executed yet another query that groups the data by the file id and sorts them by number of copies in reverse order:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/fe10d942bfb8f347a82ccdc6dd35d0e9/href">https://medium.com/media/fe10d942bfb8f347a82ccdc6dd35d0e9/href</a></iframe><p>Here are top 20 results:</p><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2Fembed%2FshrmamvWoe0o5vCAE&amp;url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2FshrmamvWoe0o5vCAE&amp;image=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fstatic.airtable.com%2Fimages%2Fsplashimages%2Fairtable_static_shot_simple.png&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;type=text%2Fhtml&amp;schema=airtable" width="800" height="533" frameborder="0" scrolling="no"><a href="https://medium.com/media/0adb3c35cf3fcde40dbd7f6b38b6101f/href">https://medium.com/media/0adb3c35cf3fcde40dbd7f6b38b6101f/href</a></iframe><p>JSON_ERROR in repo_name column means that a package.json file can’t be parsed, there are multiple reasons of that:</p><ol><li>Empty or invalid file, even best of us sometimes forget to clean up before pushing the commit.</li><li>package.json is dynamically built on compile-time, here is the example:</li></ol><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/a413bf8d973354c9cfe9cb43aa8af3f1/href">https://medium.com/media/a413bf8d973354c9cfe9cb43aa8af3f1/href</a></iframe><p>But now let’s take a look at the 5th row of the results..It says that there are 47343 instances of <a href="https://github.com/juliangruber/isarray/blob/master/package.json">package.json file of isArray repository</a>..and actually there are more than that! The total number is <strong>99692..!!WOW!!</strong></p><p>Think about it — the module that exposes very basic code:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/970e3334d5ec5d987f4df93c71d918e2/href">https://medium.com/media/970e3334d5ec5d987f4df93c71d918e2/href</a></iframe><p>was used in thousands of projects all over Github, as the direct dependency or the “inherited dependency”(project B depends on project A which depends on “isarray”). Moreover, it seems that people don’t care about keeping the source code clean and commit node_modules folder of their projects carelessly, here is the snapshot of some paths to isarray’s package.json</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*TlSJeMzRGuHatOyqPJBkbQ.png" /></figure><p>Is it just me or it resembles cancer? A ton of duplicate , sometimes sick, code spread over all the GitHub/npm ecosystem and increased the size of the projects’ source code, like a tumor. This is why Github hosts over 8 Million package.json files and only ~1.5M of them are unique.</p><p>Okay, I got it..developers are pretty busy and don’t have time to write a line of code that checks if the input value is Array..but look, even if you are going to use that module, what on Earth makes you keep the node_modules folder in the source code? Is there any valid reason of that? Please check .gitignore file of your project and make sure that it <a href="https://github.com/github/gitignore/blob/master/Node.gitignore">excludes node_modules</a> from the files to track, don’t contribute to Github Cancer development.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=180db780d99d" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Using BigQuery Github Data to find out open source software development trends]]></title>
            <link>https://medium.com/@sAbakumoff/using-bigquery-github-data-to-find-out-open-source-software-development-trends-e288a2ca3e6b?source=rss-13832a65f7f1------2</link>
            <guid isPermaLink="false">https://medium.com/p/e288a2ca3e6b</guid>
            <category><![CDATA[open-source]]></category>
            <category><![CDATA[github]]></category>
            <category><![CDATA[javascript]]></category>
            <category><![CDATA[bigquery]]></category>
            <category><![CDATA[npm]]></category>
            <dc:creator><![CDATA[Sergey Abakumoff]]></dc:creator>
            <pubDate>Fri, 12 Aug 2016 13:57:15 GMT</pubDate>
            <atom:updated>2016-08-12T13:59:17.651Z</atom:updated>
            <content:encoded><![CDATA[<p>The <a href="https://medium.com/@sAbakumoff/using-bigquery-github-data-to-rank-npm-repositories-ecf8947a1182#.k6y045vg4">previous story</a> showcased how the public data and affordable computing power could be leveraged to explore the world of open-source software : the code analyzed the <a href="https://docs.npmjs.com/files/package.json">descriptors</a> of the <a href="https://www.npmjs.com/npm/open-source">npm registry</a> entities to find the most popular ones. This article uses the same methods to infer the JavaScript development trends from the npm’s public collection of packages of reusable code.</p><p>The descriptors of npm packages are kept in files called “package.json”. Among other things, they contain the list of the keywords that “help people discover a package as it’s listed in npm search”, for example:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/68ca87aa92046cfe82a209eb67bd7360/href">https://medium.com/media/68ca87aa92046cfe82a209eb67bd7360/href</a></iframe><p>So, let’s find out the trending keywords, shall we? The <a href="https://medium.com/@sAbakumoff/using-bigquery-github-data-to-rank-npm-repositories-ecf8947a1182#.k6y045vg4">previous story</a> explained how to select the contents of package.json files from the <a href="https://cloud.google.com/bigquery/public-data/github">BigQuery Github data</a>, the results were saved to “githubdataqueries:NpmStat.package_json_content” data table.</p><p>Before plunging into the query composition and results, let’s look at the amazing <a href="https://cloud.google.com/bigquery/user-defined-functions">UDF feature</a> of the Google BigQuery that allows developers to run a custom JavaScript code within the SQL query. A user-defined-code accepts the single data row as input and produces zero or more rows as output. Selecting the keywords from the content of a package.json file begs to be implemented in JavaScript because JSON is a subset of JavaScript and it can be used in the language naturally! Here is the self-explanatory code of the user-defined-function that emits the list of the keywords of a package following by the SQL query that uses the output of that function to rank the keywords:</p><iframe src="" width="0" height="0" frameborder="0" scrolling="no"><a href="https://medium.com/media/7b76d91a77f5fd1487f6903a04a67220/href">https://medium.com/media/7b76d91a77f5fd1487f6903a04a67220/href</a></iframe><p>Can you guess trending keywords?</p><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2Fembed%2Fshr8qodEjGeJ5Iuln&amp;url=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fairtable.com%2Fshr8qodEjGeJ5Iuln&amp;image=https%3A%2F%2Fround-lake.dustinice.workers.dev%3A443%2Fhttps%2Fstatic.airtable.com%2Fimages%2Fsplashimages%2Fairtable_static_shot_simple.png&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;type=text%2Fhtml&amp;schema=airtable" width="800" height="533" frameborder="0" scrolling="no"><a href="https://medium.com/media/f3a8e784275a1c362f71839fe4555f23/href">https://medium.com/media/f3a8e784275a1c362f71839fe4555f23/href</a></iframe><p>Random thoughts on these results:</p><ol><li>Keywords relevant to text processing : “string”, “text” and “parser” are in top 10. I can guess that the root cause of it is the JavaScript <a href="https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/String">built-in String object</a> does not expose functionality that developers use on a daily basis. Fortunately npm registry has a ton of LeftPad(wink-wink) and other useful string utils!</li><li>“cli” that stands for “command line interface” is in the 2nd place, closely-related “terminal” and “console” are in 9th and 11th respectively. A possible explanation is:</li></ol><ul><li>CLI packages only make sense for node.js applications</li><li>the initial node.js release supported only Linux</li><li>CLI is at the heart of Linux &amp; Unix systems</li></ul><p>3. “http” and “browser” are in top 10. That’s simple : it’s all about web nowadays!</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=e288a2ca3e6b" width="1" height="1" alt="">]]></content:encoded>
        </item>
    </channel>
</rss>