Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Sunday, February 19, 2017

"I know what you did last summer.. and pretty much everyday!" - Check out your life catalogued on Google


Couple of weeks back, I was in Bangalore to conduct a workshop at the National Institute of Design (NID), Bangalore. After reaching the hotel, I googled the NID location on map to check distance and the commute time. As I was about to close the page, I stumbled upon a small link on the left that said 'You last visited 11 months ago'.


I was surprised, since I had indeed been to NID a year back, for an earlier workshop. Intrigued, I clicked on it and saw the below:



Now this got interesting since it showed the actual date of visit in 2015 and not just that, but where I stayed (Ivory Studio One at Indiranagar) and when I travelled that day, how long I was at NID, and when I left. And then I noticed the bars at the top which seemed to have similar catalogued information, for other days as well. To investigate further, I clicked on the next day. Now it started getting beyond interesting:



This showed my itinerary on the next day, after I had got back to Hyderabad. It had a complete chronicle of the day - the time I left home, the bakeries I had visited to book a cake for my kid's first birthday party. The cake pictures I had taken at the bakery on my mobile phone (and which was backed up on Google photos) is neatly catalogued alongside.



The timeline continued to show Karachi bakery, the last one I had visited and where I finalized the cake. Then, it showed my visit to the printers for the birthday cards - and the rest of day at office - time I travelled home - the route I took that day - and Google's guess of whether I drove, cycled, took train or walked, perhaps based on the speed of movement. Now this was just too much of my personal life catalogued, and it was quite eerie! The only apparent comfort being Google's assurance that this information is safe & private, until I choose to make it public.


Finally, to understand what else was up for grabs from my google-created life travelogue, I explored further and discovered a variety of rich information - all my trips by year; every single place I had visited; the favorite joints; average travel per day and so on and on. All this information was available for a historical period of 3 years, which was when I transitioned to an Android smartphone, and incidentally handed over my life to Google!



Now, all along I knew Google has been storing my data - my Google searches, my Android phone location, map searches, photos backed up, my office inbox, my calendar. Being a data afficionado and one who takes out time to track all possible data points in the daily events, this was a conscious choice and something I have been excited about. But what caught me by surprise here was the neat processing and linking up all of these together to create a rich and frightening catalogue, which leaves nothing to imagination.

For now, I will definitely continue giving out my data, but will be a bit more circumspect about the possibilities.

Sunday, February 12, 2017

Hans Rosling - a tribute


In the past week, the discipline of Data Visualization lost Hans Rosling, a stalwart & leading proponent of Visual story telling. He passed away at the age of 68. That he was a master story teller with data is widely known, and one can get a glimpse of some of these brilliant narrations in his TED page.

A key highlight of all his talks was to demystify statistics and make seemingly boring numbers come out alive on the screen. One of the reasons he stands out in the Visualization world is his unusual combination of diverse skillsets - depth in statistics, eye for visual design and riveting presentation with screen presence.

His efforts have been lauded for taking pains to simplify statistics for the common, non-data audience. He has gone to the extent of explaining heavy concepts like population growth and climate change through Lego blocks and other physical props. Experimenting with innovative ways of presenting data, he had also used Augmented Reality to present data in 2010, several years before AR and VR came into mainstream parlance.

Hans Rosling was a big inspiration for us, in the early days of Gramener. We have leaned on his vast repository of work as part of our talks to educate audiences and in introducing concepts to our clients. Gapminder, the non-profit website and visualization tool that he founded were used as benchmarks in our early work. Over the years, I have continued to use some of his brilliant videos in my Visualization workshops, including one just yesterday.

Thanks to his pioneering work, the industry has benefitted immensely and his legacy will be carried forward. If you're looking to get started with Hans Rosling and his body of work, I'd recommend to take a look at the AR talk from 2010 where he covers 200 years of global history, the wealth and health of nations through a 4 minute talk.



Monday, November 28, 2016

Evolution of Gramener Design Toolset


Earlier this month, I had written an article in our company blog. I'm reposting it here again:


At Gramener, we have been continuously evolving our Design process over the past years. These improvements have been to stay in tune with the emerging trends in design, adopt industry standard tools and create a custom Design framework that helps us deliver outstanding visualizations.

This post covers the ‘tools’ aspect of the design improvements we’ve implemented, and it discusses the challenges faced and considerations on coming up with a pertinent toolset to cater to Gramener’s core offering of Information Design & Data Visualization.

Earlier Process brief and toolset used:
Until a year back, the primary tool we used during the design phase was paper-pencil for creation of Design concepts, while the actual designs were created on Powerpoint. For most engagements we had a low-fidelity design as deliverable, wherein the paper sketches were translated to a basic representation on MS Powerpoint using snipped images of charts and other basic dashboard components.



In certain engagements, when there was a need to show a closer-to-actual representation, a high-fidelity design was created, again on Powerpoint using imported SVG objects or drawn chart elements. There was almost no prototyping or demonstration of interactivity, save the occasional powerpoint slide transitions. The need for an internal Design library was met by having all designs stored on the Gramener file server and exposed on a searchable, minimalistic UI, that was spruced up with basic previews and meta-tags.



Given the ability of Gramex, Gramener’s platform to quickly pull out charts from the engine’s library and setup a basic, working version of visual dashboards, historically, there was not much of a need for a standardized design tool. Hence, Powerpoint was a quick and light alternative that fit in well with the skillset of Data Consultants, which is a role comprised of functional analysts, who had innate comfort with MS Office rather than the Adobe suite of products.

Evolving needs:
With the evolution of projects done by Gramener and the rapid scale-up in clientele and team size, the need was felt for a rethink of the above mentioned stack. With a large number of first-time visualization adopters amongst clients, we sensed their comfort in reviewing solutions with a high-fidelity design that showed visual design aspects as close to the final solution as possible.

With increasing functional complexity and data size of our visual solutions, Data Consultants had to spend more time in the solution conceptualization and data analysis phases, whereas there was an increased need for additional support during the Design phase.

Challenges faced:
In summary, the key challenges faced with the above simplistic process & toolset were:
  • Variation in quality, finesse and look-and-feel of designs created on Powerpoint
  • Long cycle time for design creation, with an often cumbersome process for putting together the occasional high-fidelity versions
  • Teething challenges in development handover & translation of the design
  • Need for multiple design reviews during development phase, coupled with rework
  • Difficulty in demonstrating state transitions, interactivity and user flow within a visual application concept


Alternate Solution:
Given these challenges and the additional considerations of scalability & rapid replicability, we went about evaluating changes needed in the process, toolsets & framework. We spoke to the design community and took first-hand advice from experts in these areas. From the tools perspective, we evaluated a variety of visual design and prototyping tools including Adobe Photoshop, Illustrator, Balsamiq, Axure and Pinegrow, amongst others. Based on considerations of fitment to our visualization lifecycle, availability of complementary skillsets at Gramener and ability to address the challenges outlined, we zeroed in on the following:

Sketchapp – for Visual Design:
The vector graphics editor from Bohemian Coding has been rapidly gaining popularity and has quickly built its own community of loyal users. With addition of new role of Information Designer at Gramener, this tool helped us in the following ways:
  • We found that the tool was very easy to pickup due to its intuitive usability, perhaps closer to Powerpoint. It also had ample tutorials and a robust support ecosystem
  • By design, the tool was meant to create vector objects and naturally fit in better for dashboards and web applications, while other tools were heavily skewed towards graphic design
  • Has a thriving ecosystem of plugins for productivity improvement, and importantly provides for easy export of style sheets to aid development translation
  • Comes at a relatively economic price compared to popular options, Apple hardware prices notwithstanding




Invision – for Workflow, Prototyping and Design library:
A leading prototyping, collaboration and workflow platform used by several design houses around the world, this tool checked-off multiple items in our requirements list:
  • Provides an end-to-end design workflow solution with useful admin features
  • Has native integration with Sketch and hence it automatically syncs, imports and stores assets from Sketch files. Automatically creates style sheets & enables direct look-up
  • Supports basic prototyping needs to show interactivity and transitions
  • Has useful collaboration & commenting features, and integrates live design presentation and review capability
  • Doubles up as a repository with versioning & hence can be used as a design library




In Summary, below is the overall process that we have arrived at, which has been working well for us and has addressed most of the above-mentioned issues we faced:



Friday, September 30, 2016

Google Now and its Indian predictions

I've been tracking Google Now for a while. However, Google did not push most of the features in India, though they have been rolling out incremental capabilities in US and few other geographies over the past years. 

In the past few months I was excited to see Google Now cards being gradually pushed onto my android mobile. I've been checking the cards, customizing preferences and feeding inputs based on the suggestions it had been surfacing. For instance, there are cards for weather alert, stock movements and news story suggestions based on browsing history. Few weeks back, I was glad to notice stories suggested from my Feedly account, based on Google Now's integration with the Feedly app. 

Couple of days back, I was amazed to see Google Now push a travel card, alerting me to leave for the office at 9.10 AM. I later figured that this was based on a meeting in my calendar that was setup at my office; this was combined with my home location, and by figuring out the possible routes and travel times, Google noticed a surge in traffic on that day. Per its calculations, the usual 30 minute journey was supposed to take 50 minutes, and hence it prompted me to leave immediately. 

Amazed with this alacrity, I left home a bit earlier and indeed found the traffic to be excessively bad that day. I duly checked back with Google on the best route for the day, and it showed me a off-road route, under the Hafeezpet railway bridge as the shortest one, while my usual route was projected to take 40 minutes longer than usual. Grateful to Google, I took up this route which was something I used to take very occasionally. True to the predictions, I encountered very less traffic on the route, inspite of all other roads being clogged.

With everything going picture perfect right from the start & in a dreamy state, I was woken up with a rude shock when I saw that the underpass road was filled with water upto 3 feet, from the previous day's heavy rains! That was perhaps the reason for less traffic on this particular route. Taking a U-turn, I went back to my usual route, and spent an hour longer commuting to the office. I had the Google assistant turned off for the rest of journey.

The Indian market is indeed unique and a tough one to crack! Hope the Now Cards algorithms learn its peculiarities soon enough to get accustomed to this market.

Saturday, September 03, 2016

Facebook's misstep on data privacy with Whatsapp


Facebook received bad press yet again last week, after its internet.org fiasco in India, several months back. You might have got a notification on your Whatsapp that the 'terms and conditions' had changed and you'd need to 'accept' them to continue using the services. One of the key changes in this was the user's implicit permission to let FB use their Whatsapp profile info to sell more targeted ads on your Facebook account. 

Image source: Techcrunch
And, this created considerable outrage, with a lot of messages going viral (within whatsapp!). There were talks of privacy breaches and how Facebook has started invading more spheres of our private lives. The messages also educated users with a simple set of steps on how to turn this setting 'off' in Whatsapp.

If for a moment you take a dispassionate look at the whole thing, it doesn't appear to be so alarming. Here is a parent company (Facebook) trying to cross-leverage its presence and services with a subsidiary (Whatsapp, which it bought for a bomb of $20 Billion), to better monetize the user base. The users were anyway getting both the services for free. And the new terms clearly stated that only the whatsapp user profile details would be used and none of the chats, groups or other interests would be shared.

Then why the outrage? Its not new in the B2C space for companies to cross-leverage or cross-sell services across their spectrum of products or subsidiaries. Take Google for instance, who has systematically achieved deep integration amongst their wide gamut of products, wherein consumer intelligence from one product enriches the others. But, the fundamental difference here is that the features for the user have always come in first and hence have been well received (though not without its share of suspicion); like the smart Google Now that simplifies your life. Facebook has erred by putting the carriage before the horse, and attempted to monetize first without linking the two for user's gain.

This apart, a side issue has been the perceived dishonesty or malicious intent in the way users have seen FB roll this out. The nature of an implicit, hidden agreement that kicks in when a user 'accepts' T&C certainly didn't help. When you try checking-off the box to disagree on usage of your profile info for ads, the prompt checks 'Are you sure you want to do this? You'll never be able to change this ever again'. Why should this sound like a once-in-a-lifetime favor that FB is doing you. After all, its a user preference and one might be okay to go back and enable the link when they see some benefits coming their way.

Finally, another undercurrent for all of this is the extremely accurate and targeted nature of the ads on Facebook, which has put off a lot of people. Actually this is one area FB must be congratulated for the accuracy they've managed with their algorithms! Unfortunately, the market at large doesn't see it that way. The advances in analytics have been exponential in a brief span of time in an unregulated market, that unpeople haven't been able to fathom it yet. 

A personalized recommendation is not seen as a smart salesperson suggesting just-the-right product, but rather like an intruder who has not only got into your house without permission, but has setup a canopy right in your living room to sell stuff by overhearing what you talk to your family!


Monday, August 15, 2016

Flight ratings - whats the big deal?

Booking an early morning flight, for business, is always a tricky thing. Depending on your meeting start time (say 10 AM), you'd have to allow room for some buffer beyond the budgeted time, for landing at the airport and taking a taxi to the meeting location. You'd definitely need to allow atleast 45 to 60 minutes buffer time for any delays. 

But what if your flight gets delayed by upto an hour? All your pre-planning and scheduling goes right out of the window. So, thats some more slack time needed. However, the frustrating thing for me is that all this buffer cuts down on 'precious' early morning sleep time. Proper sleep on the day of travel is anyway a rare commodity, since I've found that you need to be up by 4 AM for a 7 AM flight. 

Same is the case with the night return flights. You budget time to get back home at an 'earthly' hour, say by 11 PM. Flight schedule changes and airport traffic congestion easily add up another 60 to 90 minutes.

This brings me to the question. Why don't travel sites show ratings of a flight's past performance? Is it usually on time? How is the service? How is the food quality? This can easily be aggregated from customers, post the travel. Since delays could also be due to airport traffic congestion or other reasons beyond a carrier's control, the feedback on timeliness could be on whether the boarding happened on time.

I'm surprised that none of the airline travel webistes do this today. This is a tried and tested feedback loop in other industries - think cab aggregators like Uber/Ola, hotel booking aggregators like Tripadvisor. In the travel segment, by aggregators like redbus (sample given here). The closest site I could find was flightstats, which is a dedicated website on past flight statistics and has other means of data collection since it doesn't have a customer feedback loop to leverage (see image below).


Simple bus ratings from redbus.in FlightStats rating summary

Saturday, February 27, 2016

Sandhill Article: Building a team to deliver Big Data's promises

A whitepaper I had written on building and scaling Data Science teams was published by the Business Strategy online magazine, Sandhill in the past week. As a coincidence, we celebrated Gramener's 6th founding anniversary this week, hence its been great timing. Here's the full article:




How to Build a Team to Deliver on Big Data’s Promises


  • author image
Big data, which has caught the fancy of people worldwide, across disciplines, seems to be maturing from the ”next big thing” to providing business value for enterprises. As a key trend shaping the market, it continues to hold sway over all stakeholders in this ecosystem, whether it is the millions looking to make a career out of it, thousands of enterprises wanting to leverage data for business gains or the rapidly mushrooming set of new and established players who intend to provide solutions in this space. However, one question that baffles the big data world is “How does one build and scale data science teams to deliver consistent business value from data?” 
Relying on the adage “experience is the best teacher,” I draw upon the experiences of building a delivery organization to provide customer value from the big data promises. Need for multi-disciplinary skills, dearth of talent, intense competition to hire talent, limited hiring dollars and a fledgling brand all made the task tougher. Hence, this article discusses the realistic action in the trenches rather than a few concepts. 
Setting up a data science team 
For enterprises aspiring to put big data to use, the broad approach to a sound solution is a three-step process:
  1. Consultative solutioning to identify the business problems, define and scope out the right perspectives
  2. Pertinent analytics to derive insights and hidden patterns from data
  3. Data visualization to bridge the last-mile disconnect by converting numbers into visuals, to present the information and insights from data. 
Let’s now look at the key challenges that an organization would face in building and scaling such a big data solution and how to address these challenges. 
1. What mix of skills can deliver value? 
Delivering a robust data science solution calls for a multi-disciplinary skillset across four broad areas:
  1. Domain skills to identify the right business challenges and come up with solutions
  2. Quantitative skills needed to apply math and statistics for extracting insights from data
  3. Design skills to present information in a creative, aesthetic and usable manner
  4. Technology skills to leverage deep programming and data technologies for scripting this end-to-end analytics and visualization solution. 
2. How to hire the right skills 
With companies struggling to hire talent with good skills in most of the above areas, it is next to impossible to get a combination of all skillsets in one person. One solution is to carve out new roles in data science along functional and technical lines by bundling a set of related skills that is closest to the solution offered. 
The intent is to hire people with complementary skills, who would come together as a multi-functional team to deliver an engagement. For a bootstrapped startup, it’s a sound strategy to start with talent in known circles and through direct referrals, wherein the initial hires can be trained on the job and supported on existing skill gaps by the senior members or founding team pitching in. 
3. How to attract the right people 
The war for talent is a perpetual problem for most companies, and new-age startups take a different approach to address this issue of hiring great talent. These companies consciously invest a good amount of time speaking and participating at relevant big data conferences, public forums and partnering with educational institutions offering data science courses. While this aids with branding and creating a buzz in the industry, it also helps get closer to qualified talent and attract them from relevant circles. 
As a side benefit, this could also have an indirect fallout on sales lead generation by helping add qualified client leads to the sales pipeline. 
Additionally, targeting the key movers in online technology forums like Github, Stackoverflow and the like can be very beneficial for companies in identifying lead and senior profiles. The focus at this stage should be to make the hiring process efficient by pre-qualifying candidates and helping preclude the inordinately long cycle times and low success rates associated with the traditional hiring process. 
4. How to train the team across multi-disciplinary skills to deliver sound solutions 
With a growing team, it’s imperative to address the problem of repeatable delivery early on by putting together a sound delivery framework with clearly defined processes, roles and responsibilities, and deliverable templates and artifacts. One must keep a continuous focus on training and upskilling the team across roles by creating or sourcing content to run internal training programs. This can be supplemented by online courses and guest lectures by experts from the industry. 
All through this stage of growth, one must retain a laser sharp focus on clients and ensure that there is consistent and considerable value-add to the business stakeholders by leveraging data to smartly solve the business challenges. 
Adjusting sails for the next wave of growth 
5. How to correct chinks in the armor that impede scale and decentralization 
As organizations scale, it is imperative to reexamine the systems and processes to identify the need to adapt to changed market needs and internal dynamics. Often when companies breach the golden team size mark of 100 employees, they run into typical scaling issues around employees, processes and the quality of client-facing solutions. 
This can be tackled through a critical review and reorganization of existing processes in order to overhaul all those practices that don’t fit well with the theme of rapid upscaling. At this stage, organizations need to be open about decentralization and empower the second line of leaders who can carry the organization forward. 
Also, the convenient and comfortable practices would have to be given up in favor of more objective and standardized processes that can be rolled out across a larger team. 
6. How to improve the skills-mix by growing breadth while also achieving depth in focus areas 
Organizations at this stage need to focus on deepening the skills and knowledge areas to improve the quality of their solutions. A relook and expansion of roles by unbundling responsibility areas to allow for deepening of skills and knowledge areas can prove beneficial. At the same time, one needs to look at broadening the portfolio with complementary offerings to provide a well-rounded solution. 
Towards this effect, Centers of Excellence (CoE) or “horizontals” can be carved out within the organization in the core areas of data science, information design and technology to help achieve the needed depth in big data skills. These horizontals can take on the mantle of expanding and providing specialized training to the teams for all up-skilling needs in the organization while also taking a lead role in the complex, specialized implementations for clients. 
7. How to adopt practices in hiring for scale 
With requirements calling for greater numbers in hiring at this phase, the channels can be expanded by signing up with selected strategic hiring partners, apart from leveraging innovative techniques in data analytics through online channels for hiring qualified talent. It is also a sound practice to run hiring hackathons and data science contests on sites like Kaggle to get the right level of attention from prospective candidates while also opening up the possibility of hiring in bigger numbers. Employee-driven referrals can start yielding fruitful results at this stage. 
8. How to move up on process and solutions maturity 
At this next stage of evolution, it is critical to deepen the relationship with the client by taking on an advisory role and hand-hold the enterprise in chalking out a comprehensive data analytics road map. The maturity levels in solutions offered moves upstream, from delivering business value in chosen areas to looking at the enterprise end to end and advising on the set of strategic initiatives and moving clients up the data leadership hierarchy. 
In conjunction with this, the organization and delivery process must be re-bolstered by focusing on scalable processes and delivery excellence to support the growth needs. By weaving in the expanded roles and responsibilities of all individuals in the organization, a robust performance review process can be established to enable continuous career focus and growth for the team. 
Summary 
Building and scaling a big data organization is a continuous and challenging process. One has to continue work methodically on the above-mentioned spectrum of areas, coupled with reviewing and reorganizing at the right intervals throughout the growth stages of the organization. 
In spite of a constantly shifting base along with the many moving parts within and outside the organization, it is critical to retain an open mind-set and nurture the core ethos that is the lifeblood of an open, startup organization: innovation, technology-focus, client-centricity combined with an open culture and fun at work. By retaining the focus on these critical growth factors and addressing the scaling challenges, this cycle of conceiving, scaling and maturing of an analytics delivery organization can indeed be made a reality. 
We continue to unlearn, relearn and scale to the next level in our journey towards becoming a mature big data product and solution provider. As more and more customers commence their nascent big data journeys or look at moving up the maturity value chain, organizations like us that offer solutions in this space have a critical responsibility to not just merely aid them, but transform constantly to deliver concrete and lasting return on investment throughout the journey. 
Ganes Kesari is the VP of products and consulting at Gramener, a data visualization and analytics company. He tweets from @kesaritweets and can be reached at ganes.kesari@gramener.com