Mean, Median, and Moose

Mean, Median, and Moose

By
Download on the App Store

Mean, Median, and Moose episodes

  • Food and Agriculture Data in Canada

    This week on Mean, Median and Moose we’ve got data on food, food processing, and eating out!

    Garden Planning Tools

    Frazier Fathers

    So I decided to dig into something a little different this month, I have a vegetable garden in my backyard that I have always just sort of “winged” in terms of planting since moving into my house. I have torn out a garden that I should have known was going to be unproductive (too much shade and too close to the fence) and built a new semi-raised bed out of old deck boards (mainly to not have to pay to get the wood hauled away). 

    This year, I am thinking I should actually put some effort into planning my garden so I went looking for online planning tools to see if there was something user-friendly and helpful for my gardening adventures. This led to much googling and asking ChatGPT about potentially available tools, then to Reddit and down rabbit holes weeding through debate as different people offered up different software or tools. I have selected a handful of options to sign up for an a few other poor experiments

    My first test was with the classic Microsoft Paint! It ended with just a bunch of squares and shapes on a screen that really didn’t add value.

     

    Online one of the first places that I found success for a general garden plan was draw.io. This isn’t actually a garden tool but rather a diagram and flowchart tool. That being said, if you know the measurements of your yard, the grid it lays out shapes on can easily be used to ensure proper scale of the bed and other features in the yard. 

    The tool allows labeling of shapes but is not great at allowing you to put shapes inside shapes and has no – plant specific shapes that I could easily find. I experimented with adding lots of little green dots similar to the evergreens illustrated above but they got really messy and there was as the snapping feature kept trying to connect them with lines or to edges. 

    Garden Planner 3 is a highly rated tool on Reddit and has a 15-free trial which I played around with. Designed to plan gardens and garden beds as the name suggest it is a far more bespoke tool for your backyard garden planning! There are literally hundreds of symbols and models for the a range of garden plants but also other backyard features like paths, ponds, gazebos, decks, bbqs and much more. 

    Each of the features can be assigned a specific size in both imperial and metric scales. So if you know your garden bed is 5 feet by 10 feet, you can create a bed space that is that size on the screen. Within the “Design Garden Bed” feature you can actually layout specific beds with symbols sized to the approximate spacing requirements in a bed. 

    In the image to the right you see a 10 foot by 4 foot garden bed. The tool represents Tomato plants as 3×3 structures at full size while pepper plants are represented as a 2×2 shape. 

    One of the problems with the tool is when trying to put non-straight lines. As you can see by the walking path in my backyard the weird shape it has 

    Growveg Garden Planner – once again offers a free 7-day trail based on annual subscription of $60. I found this through the Youtube Channel with almost 890,000 subscribers. The online tool with accompanying apps, this one is far more robust beyond just planning your garden layouts; it also gives you features and materials that support you through the entire growing season – including schedule planners, weather and watering tracking, a journal tool as well as live chats, group discussions and email reminder. 

    When you are setting up the app, they ask for your address which connects to the nearest weather station to align your garden schedule with traditional frost dates as well as weather right in your app – – to help you decide if you need to water! Based on this geography, a calendar for seeding indoors, seeding/planting outdoors, growing periods and harvesting. 

    The app includes a good seven step training program which the image below is part of their tutorial garden plan. Each icon is visualized at its proper spacing for proper growth and elements like trellises and stands can be added for climbing plants like beans and cucumbers. You can click on them in the app and it will bring you to the journal and tag the plant for a reminder. 

    The journal is actually a pretty powerful feature as you can take pictures and tag them use the data year over year (I suspect). Within the journal it also contains data about what companion plants, pests, and potential solutions to those problems.

    I did get a little lazy and did not build my backyard in the tool (see generic garden above). It has fewer generic yard and lawn icons, bells and whistles than the Garden Planner 3 but it is better overall at various garden elements and outcomes. Another minor quibble is that for some items that plant very closely together, the graphics become almost impossible to see (Garlic above). Certainly the close planting is viable and optimal but if I can’t count at a glance how many seeds/bulbs I need it makes it tough to use.

    The Verdict

    The problem with all of these tools is that in many ways just like learning any new app, there is a learning curve. I am recording this in February, I am probably going to start seedling things in the coming weeks in my kitchen next to my window. To date I have probably invested an hour or two in each of the tools poking around and trying things. If I just planted a few extra seeds in that time the outcome of sub-optimal placement/planting in a month from now would probably be offset. 

    The second issue is that all of these options still require a core understanding of your yard, garden and how it will grow. For example, just outside of my yard my neighbour has a tree that blocks direct sun in certain parts of the day in certain parts of the gardens. The various diagrams illustrated above do not take into account sun, soil or other elements that are site specific. 

    Overall all of these tools take away from a key thing for me…. Not being on my phone/computer! So for that reason, just plant some seeds and unless you are struggling for survival through your garden produce you are probably fine with sub-optimal plant and garden placement. 

    If you are willing to invest 5+ hours mapping your actual or potential gardens upfront then the tools explored certainly could find value over time but the opportunity cost on an exist backyard or something could be a barrier to enter.

    Little Foot Foods

    Rachael Myers

    Hi! My name is Rachael Myers. I co-own Little Foot Foods. A Gourmet Pierogi Manufacturing Company located in Windsor Ontario. We have been in business for a little over eleven years and I am excited to explore our growth with you through the world of statistics! 

    It all started in 2013 at a local community centre kitchen. With a beloved family recipe, a hand washing sink, commercial refrigerator – that only had one shelf in it, three chest/deep freezers for storage and our mvp! – a commercial dishwasher. Each bag had 1 dozen handmade pierogi in it.

    Looking back at the data that I have, it’s pretty sketchy from this time. In the beginning we only had enough production capacity to give away free samples at events. No sales! We requested people put in orders for delivery….. and they did!  There were definitely cash sales during this time that were not captured in our Point of Sale (POS) system as we got up and started. 

    Because we did not have a store, we offered free delivery two days a week. During this time we sold our pierogi for $9/dozen and went from 1 flavour the OG – Potato, Cheddar and Onion to a dozen flavours practically overnight. All our sales were generated at local events and farmers’ markets until we got our first wholesale contract.

    Each week we processed 60 lb yellow and 60 lb of russet = 120lb total of potatoes and  7kg cheddar cheese! We had 3 local retailers – Dressed by an Olive, Robbie’s Gourmet Sausage Co. Lee and Maria’s and 8 varieties of flavors – OG Onion Cheddar, Classic Ceddar, Bacon Cheddar Onion, Spinach Feta Ricotta, Jalapeno, Enchilada, Chorizo Cheddar, Balsamic Mushroom Onion. The last two are an example of collaborations – Robbie’s Gourmet Sausage Co and Dressed By an Olive. 

    We handmade pierogi for 3 years out of this space. By the time we outgrew it we had one full-time employee and amazing potato peeling skills! At this point we had to make some tough decisions. The business was unsustainable as hand made pierogi are incredibly labour intensive and our main competition is churches who utilize volunteer labour.

    The decision was made when a friend of ours found a pierogi making machine in Kijiji and we decided to take the leap into a full time endeavor. Moving into our next growth phase we mechanized our production to some extent. An electric potato peeler which cleans and peels in addition to the pierogi making machine that I described above. This revolutionized our output.

    We increased our production output – running production 2 days a week! Having room to freeze 2500 individual pierogi or 179 bags (which are now packaged by the pound/454g) of pierogi each production day. We invested in retail packaging – fancy picture, appropriate nutritional labeling the whole enchilada

    During our time on Tecumseh rd. we used.160 pounds of yellow and 160 pounds of russet potatoes/week = 320lbs/week total potatoes and Cheddar Cheese – 17.28kg/week

    During this time we made approximately 50 different flavours of pierogi – growing collaborations, with more varieties using Robbie’s sausage, Butchers on the Block – pulled pork, Schwab’s smoked pork rib-rogi (which may be returning soon due to high demand)

    After rapid growth, we plateaued – we had run out of freezer storage space utilizing 13 chest/deep freezers.

    In the Fall of 2021, we recognized that we had maximized the output of our current facility on Tecumseh road and again we were faced with the decision of closing or expanding. Our vision of a warehouse, retrofitted to manufacture pierogi and also provide us with additional space to expand our offerings. Turns out a self-taught manufacturer has very few skills that translate into knowing how to build a production facility….. It took forever and was much more expensive than anticipated.

    We keep an ever-expanding list of flavours including Reuben, Ccorned Beef and Cabbage, and even more collaborations including one with every Farmer’s Market food vendor this past summer. Highlights being an Indian curry, spaghetti and meatballs and a Ghana spiced veggie. We have seen many food tends come and go throughout our manufacturing experience. For example, Vegan food had a real heyday in the years just preceding the pandemic. For the most part we try to stay current with trends but really it usually comes down to sticking to our root question “can we put this into a pierogi?” and usually the answer is yes!

    Supply and Disposition of Food in Canada

    Doug Sartori

    This month I’m highlighting Statistics Canada table 32-10-0053-01 “Supply and disposition of food in Canada.” It’s sourced from a variety of statistical products like the Monthly Miller’s Survey, the Fruits and Vegetables Survey which we’ve looked at before, and the Monthly Dairy Factory Production and Stocks Survey.

    The table aims to provide a comprehensive data set for the acquisition and dispersal of food products. The data set starts in 1960 and provides an annual record of the total supply and disposition of each food product, along with values broken down into beginning and ending stocks, production, imports, domestic consumption (“domestic disappearance”), exports, manufacturing and waste. 

    What I love about this data set is the attempt at being comprehensive, and the high-level view it gives of agriculture and the economy. I put together a simple visualization tool to show the data by commodity. The unit of measure is either thousands of tonnes or thousands of kilolitres, depending on whether the commodity is dry or liquid.

    Every data series tells a story. Take a look at this chart of wheat flour production in Canada. Capacity has increased, but not enough to keep pace with demand. In 1960 Canada exported something like a third of its wheat flour production. That’s dwindled over time.

    The data series for milkshakes tells an interesting story about the popularity of this summertime treat. Milkshake production is down by a factor of ten since the turn of the century. Why did Canadian milkshake demand peak in the 80s? 

    Some series show the challenges of trying to create a comprehensive data set from disparate parts. It’s clear that some series end in the mid-90s and are replaced by a different one. You can’t really make sense of these over the total period, but they can still be useful for more limited ranges. Salad oils is a good example.

    Sometimes a data series just ends. I assure you there are still dry beans being produced and sold in Canada, but they must be tracked under some other commodity name after 1992.

    Even with those caveats and limitations, this data set offers a ton of value. If you want a bird’s eye view of food in Canada, this is the place to start. 

    Are Canadians Eating Out More?

    Rashmi Krishnamohan

    Hello Hello! Glad to be back on the show! I went with the topic “Are Canadians Eating Out More?” Let’s be real—when was the last time you cooked a full meal at home? Not just reheating leftovers, but really cooking from scratch? For me, it’s been a minute! But if you’re anything like me, you probably ordered takeout not too long ago. Why is that? Are people just busy, or is there something else going on?

    So I dug into Statscan’s Canadian food expenditure data.

    The first thing I looked at was how much Canadians are spending on food –specifically, between restaurant meals and groceries. And here’s what stood out: 

    Canadians spend more on restaurant meals than on almost every grocery category. That means that on average, they are spending nearly $2,000 a year just on eating out—more than on meat, dairy, bakery products, combined.

    Even within restaurant spending, full meals take up way more of our budgets than snacks and beverages—so it’s not just about grabbing a coffee or a muffin. Canadians are replacing meals with restaurant food.

    The second biggest category is non-alcoholic beverages and other food products, like your daily coffee, pops, and sodas, coming in at an average of $1,539.60 a year. If you thought your daily Starbucks was just a little treat, turns out, it’s an investment (Oopsie :P)!

    But there’s another big expense – meat. People spend a lot on beef, chicken, and other meats. In Newfoundland, people are spending $1468.50 a year on meat, which is significantly higher than most other provinces; in Alberta, about $1,309.90 and even in Saskatchewan, it’s $1,251.70.

    As you can see below, Canadians spend most money on beverages, meat, and dairy. These are the big spenders in the grocery game.

    When I looked at the Food Expense data trend from 2010-2021, restaurant spending has always been on top. But in 2019 – there was a major surge in restaurant spending, and then COVID hit, and we all had to cook at home (for a while).

    Now, let’s talk about why this is happening. I pulled the data from the Statscan’s Food Price Index dataset for the past decade to see how food/ grocery prices have changed, and no doubt – they’ve been increasing exponentially, especially on things like meat and dairy. I know you’ve noticed it — every trip to the grocery store feels like it costs way more than it used to.

    Let’s look at meat, for example. In 2014, meat prices were sitting at an index of 144.5. By 2024, that number jumped to 208.73 (that’s a huge increase). We’re talking about over a 44% increase in the last 10 years!

    And it’s not just meat – Dairy products went from 136.45 in 2014 to 175.99 in 2024, a solid 29% rise. Even bakery and cereal products, which were once a little cheaper, have gone from 153.5 in 2014 to 200.76 in 2024.

    When I compared the restaurant meal expenditure with the Food Price Index, the relationship between these two was almost linear. In other words, as grocery prices went up, restaurant spending also seemed to increase. As grocery prices go up, people are eating out more. The logic is simple: if groceries are expensive, dining out can feel like a better deal. Think about it – $15 for a meal at a restaurant doesn’t sound so bad when a $50 grocery trip barely gets you enough for a week.

    So, I ran a predictive model for this scenario!

    I ran a simple linear regression model. The goal is to predict what would happen if grocery prices went up by 10% in the next year – would Canadians spend even more on eating out?

    I increased the Index of grocery product categories by 1.10 and my model predicted that as grocery prices go up, restaurant meals might become the more hassle-free option for many Canadians. It’s not like people have stopped cooking at home, but the rising prices are making us more likely to eat out more.

    51 min
  • Christmas Data Special 2024

    It’s that time of year again, we we try scratch up some open data on the Canadian Christmas experience! Some good ones this year, along with a special guest from Niagara college.

    Christmas Ice

    By: Doug Sartori

    The holiday season in Canada is a time for family, feasting and fun. Many Canadian children wait expectantly for Santa Claus to leave his North Pole workshop on December 24th and criss-cross the world delivering toys to all of the good boys and girls. Despite NORAD propaganda, all good Canadians know that sleighs are exclusively surface vehicles, even the reindeer-propelled kind.

    With this in mind, I considered Santa’s likely overland route to Canada. A glance at an Arctic map from the CIA World Fact Book shows that the obvious starting point for Santa’s Canadian travels is the town of Alert on Ellesemere Island, the world’s northernmost town.

    To understand the risks to Santa’s overland route, I looked at data on sea ice in Canada’s Eastern Arctic (which includes the stretch north of Ellesemere which Santa certainly must cross). The Canadian Ice Service produces regional shapefiles covering spatial trends in ice extent. They also publish these charts in GIF form, which look like this:

    These shapefiles encode data in a format called SIGRID-3. SIGRID-3 is a vector format designed by the Ice Charting Working Group, who are a coalition of the world’s ice research centers. The format was developed in 1981 and formalized in 1989, with SIGRID-3 being the third revision. The shapefile’s data table must have 17 mandatory fields along with 38 optional fields. These data fields describe the nature and extent of various ice types in the geography indicated. For our purposes, we only need the field “CT” which indicates the total concentration of ice. To calculate area, I used the shapefile’s geometry, reprojected to the Canadian Lambert conformal conic map projection for accuracy and consistency.

    There were some anomalous ice years in the early part of this century, but for the most part the extent of ice in December in the Eastern Arctic continues to be maximal. Looks like Santa can do a 72-month payment plan on that new electric sleigh without worrying that he’ll end up going for a swim on his way to Canada.

    That’s the Canadian avenue of approach taken care of, but what about the rest of the Arctic circle? For that, I accessed data from the National Snow and Ice Data Center (NSIDC) FTP archive, which provides monthly records going back to 1979. This data is a simple CSV and doesn’t require much fancy work to visualize, so I added a linear regression for a simple straight-line projection, as well as a polynomial regression for a more nuanced projection, out to 2050. This analysis provides mathematical approximations of Arctic ice loss using regression models but does not account for physical climate processes like ocean currents, feedback loops, or emissions scenarios, which are included in more sophisticated climate models. There is a data anomaly in the late ‘80s which is indicated on the chart. This chart is a little less rosy and perhaps Santa will need to consider other transportation options over the next 25 years or so.

    Python is great for this sort of work. Tools used in the analysis included:

    • pandas for data cleaning and processing.
    • geopandas for spatial data extraction and area calculations.
    • matplotlib for visualization of both time series and spatial data trends.
    • scikit-learn for linear and polynomial regression modeling.
    • You can find the source code and instructions to generate these charts on Github.

      Student Mean Median Moose Segment

      By: Lee Doucet

      I’m a professor at Niagara College in their Business Analytics program, focusing my teaching mostly on Data Visualization and Communication, having designed and created the course for the college. As a practitioner of open data, I embed open data in my course to move students beyond textbook data and let them see real challenges with working with open data. For instance, last semester, Niagara College partnered with Living Lakes Canada and learnt firsthand how important it is for adequate funding to ensure there is consistency in data collection year over year for accurate reporting. This is something you can’t see with data that is already cleaned and transformed.

      This semester I wanted to focus on housing as the cost of living is very topical. I want students to be able to investigate complex topics and produce data driven insights. When I then observed the Mean, Medium, and Moose’s methodology which emphasizes reproducibility when breaking down complex topics to present them in an accessible and approachable way, a light bulb went off. This is a highly valuable analytical skill, I want the students to be able to demonstrate their knowledge while being able to communicate it to someone on their team, such as a colleague or a supervisor, who may not have the same technical abilities to work with data.

      Often, I worry that society puts too much emphasis on the results of people’s work and not the process.  I see this through busy people flipping through to the policy recommendations, skimming the report, or just reading an executive summary only. To me, the process itself is an opportunity to engage people with how you did the work and understand if you made the correct decisions and pivots based what information was available to you. Not only that, once students correctly understand how workflows operate, they can begin to transfer those skills to their teams in the Workplace.

      All of this is lead to my interest in working with Doug from Mean, Medium, and Moose to give the students an opportunity to try something different. Housing is not an easy topic at all by any means, and over the course of several weeks, students with busy lives and jammed semesters, got together in 8 teams and produced some great reports with a supplemental dashboard on housing. They gathered data from a multitude of sources, including StatsCAN and the Canadian Mortgage and Housing Corporation, all of which are available to anyone who wants to explore Canadian Housing data. I learnt new things while reading the reports which to me is a hallmark of success. I also saw things I expected which is the struggles of affordability. In reference to that, the winning group called their project “Housing, Hurdles, and Hope”, which if you think about it, how many Canadians unable to afford the hurdles of acquiring a house in this market are relying on hope as their strategy?

      Here’s a link to the report.

      NFB Christmas short film color profiles

      By: John Haldeman

      The National Film Board of Canada is a federal government organization that funds, produces and distributes Canadian films. They produce mostly short films and documentaries related to Canada or by Canadian filmmakers. Almost all of their films are free to view on their website. Keeping with my recent theme of attempting to explore ways to visualize unstructured data, I wanted to see if I could use the films themselves as a source of data. What I ended up doing is creating visualizations representing a sort of color profile for each film taking inspiration from a project called “A viz of Ice and Fire” which was a master’s degree project in Georgia Tech’s CS 7450 course. Those students did much more than just extract the predominant colors in frames captured from Game of Thrones, but that is the part I copied for this analysis.

      Here’s the result for twelve selected Christmas or winter themed NFB short films:

      Many of these films are provoking slices of Canadiana. “The Nativity Cycle” is a film of a play put on by elementary school children in the 1950’s exploring dimensions of the story of Christ’s birth with surprising complexity and high production values. “Teach Me to Dance” is a Christmas story involving Albertan Ukrainians in 1919. “Christmas at Moose Factory” is a film entirely made up of Cree children’s drawings from a residential school in Moose Factory on the Hudson Bay. It explores life in Ontario’s North through the eyes of children in 1971 during Christmas. It’s quite an array of what is now a collection of thousands of films.

      Just for fun, here are my favourite NFB short films and their color profiles:

      The most striking is “The Cat Came Back” which careens through different landscapes while the protagonist desperately attempts to flee his feline pursuer.

      The home of Reindeer in Canada

      By: Andy Dyck

      Rudolph the red-nosed Reindeer back a part of Christmas holiday lore with an appearance in a booklet written by Robert L. May for holiday booklets to be distributed to customers of the Montgomery Ward department store in Chicago. Now, what I don’t know about reindeer could fill volumes, including the fact that reindeer and caribou are indeed refer to the same species. That said, I’m certain that with Rudolph hitting 85 years since his introduction, Santa might need to start looking for a more youthful lead for his team. In order to help out our friend Kris Kringle, I embarked on a journey to find the habitat in Canada where he’d be most likely to find a replacement. 

      I started my search for caribou in the most logical place – the Government of Canada Open Data Portal. I found 63 provincial records and 25 federal records to choose from. In order to get a national view, I narrowed my search to only federal records and further filtered down the list to only those records that would have spatial data including GDB, CSV, SHP, and XML. I found this dataset concerning the range of species at risk in Canada that includes the habitat range of Reindeer (Caribou) in Canada. 

      The format of this dataset is in GDB and I used the {sf} package in R to quickly read this dataset before doing some further analysis. I honestly can’t say enough good things about the {sf} package for R. Like geopandas for python, this package not only handles the complexity of reading/writing geographic datasets, but it also makes geographic transformations and calculations so easy that one can really focus on the analysis with needing to get too deep into the details or complexity. The dataset was fairly clean and the only thing it needed in order to make a nice looking map was a base map of Canada’s boundaries in the background and I grabbed that quickly using the {rnaturalearth} package for R and this is the result.

      I calculated the total area that reindeer (caribou) range covered and compared that to the total area of Canada in order to make a catchy title for the map that highlights that over 80% of Canada’s landmass is considered territory for Santa’s sleigh team. This is great news for Santa as he shouldn’t have too much trouble finding more reindeer in Canada.

      Out of curiosity, I wanted to know which province in Canada would have the largest share of it’s area covered by reindeer habitat. Nova Scotia, New Brunswick, and Prince Edward Island look to be out of the running from the start, but for the other provinces, I’d need to calculate the intersection of each province and the reindeer ranges first and then do the calculation of the percent of each province covered by reindeer habitat. This analysis produced the following bar chart showing us that at 56.7%, Manitoba has the largest share of its territory covered by reindeer habitat.

      Bottom line: 

      • I continue to find the combination of the {sf} and {ggplot2} package in R to be great to quick spatial analysis and plotting. I’d love to try doing this analysis using geopandas at some point.
      • The range of Reindeer (Caribou) in Canada is absolutely huge. I’d love to learn and understand more about the species and populations in Canada and how this range does or does not overlap with Canada’s human population.
      • Check out and repeat this analysis by following along with code in the GitHub repository.
      • 46 min
      • Live at “Growing your Workforce”

        This episode is special as it’s our first live recording! Time was limited so you may hear us speak quickly and have less banter, but we hope you enjoy the episode anyway.

        Job post vector embeddings (John)

        This is a small exploration of vector embedding. These are components, primarily related to the input for training and using LLMs. They also have a lot of other interesting use cases in the area of search.

        We’re going to use some job posting data provided to use by the hosts of “Grow Your Workforce”, the good people at “Workforce Windsor Essex”. We got the raw data that drove their job board and used some scripts to scrape and extract the job posting data from 1,200 job postings in the Windsor-Essex area.

        Next we translated this into a vector embedding, which is a 2,048 number description of the words and their positions for the job posting. In this way, the job posting becomes a point in 2,048 dimensional space. We used a well known LLM model to do this, but there are likely more appropriate models for this use case.

        This is Mean, Median, and Moose which means it doesn’t exist if it’s not in a chart. Unfortunately, data in 2,048 space is very difficult to visualize. Luckily there are techniques to convert this large dimensional data into a smaller number of dimensions. One such technique is called t-SNE. We used to to convert our vector embeddings to 2D space:

        Now we can chart it:

        You can see that the data is clumped together in clusters. We set it up so that you can hover over a point to understand what job posting is at that point. When you do so, you’ll see that related jobs are close together. You can do this yourself at the interactive version of the chart here.

        You can compare this “automatic clustering” by the model with the manual NOC 2015 code put into the data set by WFWE in this diagram below:

        You can see the clusters of jobs do follow similarly colored clusters of jobs. You can imagine how one might compute the vector embedding of a search term to find the closest jobs to that term. Or compute the vector embedding of one job to find similar jobs. Or compute the vector embedding of a resume to see job postings that contain similar lists of skills. All of this is a kind of “semantic search” and vectors are a common way to implement them. Add int the rest of the transformer to generate text based on the vector embeddings and you have an LLM.

        The use of emojis in workplace communication (Katie)

        So I used to work at Workforce WindsorEssex, and I mulled over updating one of my old reports, but I figured, why not explore something we all see everyday but probably haven’t seen much data on? Emojis at work. If you’ve emailed or messaged with me, you know I’m a huge fan of using emojis at work.

        The earliest Canadian data I could find that specifically focused on using emojis at work was from 2016 when OfficeTeam, a Robert Half Canadian staffing firm, conducted a survey on using emojis at work, garnering responses from more than 300 senior managers at Canadian companies and more than 400 Canadian workers employed in office environments. They asked respondents to choose 1 of 4 statements they most identified with regarding emoji use in work communications. The result? A resounding thumbs down from the senior managers with 79% considering them unprofessional, but more of a shrug from the workers, with 36% using them often or in casual communications. Can we assume most senior managers were Gen X or Boomers, while more workers were Millennials? Probably, and you’ll see why this is important in just a minute.

        Some of the more comprehensive data on emoji use in the workplace I found was American data from September 2024 when Mailsuite, an email tracking tool company, worked with the research consultancy Censuswide to survey 2013 American office workers about how they use and perceive emojis in work emails.

        Mailsuite and Censuswide were kind enough to create beautiful data visualizations so I didn’t have to. Of 9 perceptions of emoji use, the top 4 respondent perceptions were positive, with email emoji use associated with being more friendly, having more personality, being more fun, and being more approachable. All 5 negative perceptions of emoji use, including being unprofessional, annoying, cringey, less competent, and less intelligent, ranked significantly lower.

        Importantly, does why we use emojis match the increasingly positive perception of use? On the highly positive side for workplace culture, 56% of respondents incorporate emojis into emails to be friendly, 41% use emojis to soften communication, and 31% include them to show they are approachable. Then of course, 19% have succumbed to peer pressure and use them because other colleagues do, and there’s another 6% who are just being a little passive-aggressive.

        Of course, Mailsuite also confirmed what we all could’ve guessed – emoji acceptability in work emails varies greatly by generation and situation, with Gen Z and Millennials by far leading the generation pack compared to Gen X and Boomers in finding it acceptable to use emojis in work emails to colleagues, customers, and managers and even in some stickier work scenarios like telling a colleague they’ve done something wrong.

        So, what does more recent, though slightly less comprehensive, Canadian data say? The University of Ottawa’s School of Psychology seems particularly happy to support emoji research, with several studies published between 2019 and 2024. A 2021 UOttawa study on general emoji use found emojis convey information about the sender’s affect, senders that use positive emojis are perceived as being warmer, and emojis enhance comprehension of messages, with the study concluding support for the use of emojis to improve communication, express feelings, and make a positive impression during socially-driven digital interactions. 

        Olivier Langlois’ 2019 study that asked about emoji use in professional contexts found “positive” emojis were both acceptable and welcome in the workplace after some familiarity was established, while using “negative” emojis causes discomfort. Like the American study, he also found older people are significantly more resistant to and put off by emojis than young Millennials and Gen Z. Most interestingly, he notes, “This desire to break formality by using these digital images is found in the response of a 17-year-old participant who adds emojis to ‘lighten the tone of the conversation.’ A 21-year-old participant also wants to avoid frigidity by using emojis ‘to make the conversation more relaxed.’” UOttawa student Megan Leblanc’s 2023 study further confirmed the magnifying effect of emojis, finding that emojis “magnified both perceptions of agreeableness and perceptions of emotional state. Social context had no effect.” – meaning this finding applies in professional contexts too.

        Another UOttawa Study from 2024, which was the first comprehensive investigation into intergenerational emoji use in adults over 60-years-old, found that older users are less likely to use emojis, use fewer emojis, and feel less comfortable in their ability to interpret emojis. The study suggests that it is important to promote the usage of emojis across all ages and specifically older adults as it could help facilitate intergenerational interactions.

        So, based on all the data we just reviewed, I hope you can guess why you should care about emoji usage at work? Those who use emojis at work are, for the most part, and as long as they’re using positive emojis, trying to spread positivity and build strong work relationships, and this just happens to speed up workplace communications and enhance their effectiveness too. Most importantly, using emojis is a way to improve intergenerational text-based communication, which we all know has become incredibly important in this digital and multigenerational workplace! So fear not, Boomers and Gen X, embrace those emojis, and you might just find the younger generations embracing you a little more too.

        Income Inequality and Market Income (Frazier)

        Once again I return to income inequality data but as this is a workforce summit I am focusing on a specific type. Income inequality is tracked by the Gini-Coefficient/Index with a score of 0 being perfectly equal and all income being spread out evenly across a population while a score of 1 is perfectly unequal – one person holding all of the income. 

        In Canada we track income inequality in three forms – Total Income, After Tax Income and Market Income. Each of those forms are available in adjusted or unadjusted forms which group incomes by household (adjusted) or individually (unadjusted). For this I am largely talking about adjusted market incomes.

        Using data from the T1 Taxfiler database we can see how Canada and Ontario have fared over time. 

        Nationally things have stayed about the same from a market income perspective whereas Ontario tends to be a few points above the national trend. To be honest, this is largely due to Toronto. When you put the other forms of inequality onto a chart you see similar trends overall at slightly lower levels – high 0.3 as government transfers offset some of the inequality in market incomes. 

        Now part of the problem with inequality data is that outside of the Census it is limited in scope and availability. But as I mentioned on the podcast before I have a specialized dataset that covers the last two censuses from Ontario! 

        Looking at a Census division level (region) we see the following trends. 

        I find the difference between adjusted and unadjusted income as it shows how accounting for households as a unit vs individuals within the household separate as interesting. First and foremost across Ontario inequality in most regions went up between 2015-2020. 

        Dufferin, Waterloo and Niagara saw the largest increase in market income inequality increased the largest – likely having to do with some overflows from Toronto. In contrast, Lambton, Hamilton and Elgin saw their inequality in market income decline! Hamilton is the surprising one as I would assume Toronto overflows would impact it as well. 

        The difference between adjusted and unadjusted inequality is worth about 0.07 increase in inequality across these communities with the largest increases in Halton and Dufferin region and the smallest increases in Chatham-Kent and London-Middlesex. What you are seeing there is the gender and inter household income inequality play out.

        Although deeper analysis is likely needed why a particular DA became more or less equal at a glance, green areas saw income inequality decline, whereas the orange shaded saw income inequality increase. 

        Of the 6,904 Dissemination Areas (population ranges generally between 500-800 people) 4,099 saw the market income inequality increase – approximately 60% of DA representing approximately 2.7 million people – a population greater than all of the Atlantic Provinces. 

        The interesting piece of this is that these maps compare pretty strongly to some historical data. These maps look alot like inequality data from rust belt US in 2010-14 period – the lead up to the first Trump Presidency. 

        Fortunately when we look at data from After-Tax Inequality we see better outcomes, but that data includes 1 time COVID supports which have now disappeared and if you recall from the first chart from a provincial level inequality has dipped a little bit since 2020, but is still up. Market incomes have likely adjusted as well as some people were laid off during the Census etc. 

        As someone who does work in community the feeling of being left behind, is palpable now. 

        So for our workforce planning people in the room, the question is, is a job a job anymore. Is a salary a salary? What is your role in ensuring that work is good paying, and has opportunity in your community and is not just driving inequality. Something to think about. 

        Job Benefit Extraction (Doug)

        I used the same jobs data as John. Thanks, John for cleaning up all that data – it was really nice to work with.

        In looking around at recent research in this space, I found a paper published in a journal called Applied and Computational Engineering by Jie Zhou and others. This paper demonstrates a method to use Large Language Models to extract structured data from free-text job postings. Zhou reported pretty impressive results, with a refined LLM prompt producing more accurate results than a more traditional dictionary method for data extraction.

        I wanted to see if I could get a similar level of accuracy with the Workforce job data, and I decided to focus on benefits, as a proxy for job quality.

        Here’s a sample record from the Workforce jobs data. It’s really nice, with lots of structured data for analysis. One thing that’s missing from the structured data is information about benefits, but employers often provide information about benefits in the text of the job posting. I hoped it was possible to enhance this data set with some information about benefits. 

        Zhou used a model called ChatGLM3 to process a data set of around 32 thousand jobs from a Chinese jobs database.

        I wanted to compare different models’ performance, so I used Nebius, a company that provides a variety of services for people working with artificial intelligence applications. They’ve got an API for interacting with LLMs, and they offer a broad variety of models to interact with. I found the comparison tool in the API playground invaluable for comparing model performance.

        Different models have different capabilities, and are tuned for different functions. I worked through most of the list and found, unsurprisingly, that the flagship Meta open source model Meta-Llama-3.1-405B-instruct, offered the most reliable performance.

        That name is a mouthful – breaking it down, I used Meta’s Llama 3.1 model, the version released in July of this year which has 405 billion parameters. The higher the number of parameters a model can adjust in training, the more complex and nuanced performance it is capable of. More parameters can also make a model slower and more costly to use.

        In my testing, I found that most other models on offer had an error rate so high that it was observable in interactive tests. With more extensive prompt tuning, it is probably possible to use a cheaper model for this task.

        It didn’t take long before I had a prompt that would reliably return an array of benefit strings for processing. I tried a lot of different versions of this prompt, but this was one of the last ones before I abandoned it. I choose this one because it shows the way I was trying to prompt the LLM to return more structured data. A persistent problem was that it would not reliably categorize benefits based on this prompt, preferring to spit out strings pulled directly from the job posting.

         Unfortunately, I found that extracting benefit strings alone was too unreliable for a couple of reasons. There are a lot of questionable perks – low- or no-value items describing work culture or attitude – identified in job postings. I didn’t want to measure the perkiness level of job postings so I needed to find another solution. In addition, there was too much variety and at the same time too much similarity in benefit strings. I used k-means analysis to get a sense of the data, but this structure was pretty resistant to analysis.

        Eventually, I came up with a prompt that would return usable data.

        The main breakthrough was in providing a list of categories of benefit. This hint was what the machine needed. The data structure it returns is highly useful as it provides a listing of the benefit strings under each category, so we get the raw data plus the categorical assignments. This allows for a pretty solid evaluation of the quality of LLM output, and a few different paths for analysis.

        Remember those cheery but non-material  “benefits” I mentioned earlier? That’s what the “other” category is for. It’s the garbage can for dodgy benefits – The LLM rolls up most of that stuff into the “other” category, which I then throw away before analysis.

        The model I used costs $1 per thousand input tokens, and $3 per thousand output tokens. Tokens are the basic building block of LLM analysis, typically representing about four characters of text. All-in I spent about $15 on machine time to try out this method. The cost of a single run on the whole body of job postings is a little less than $4.

        Combining the data returned with the original job posting gives us a rich data set with plenty of features to use in analysis. A lot more can be done to refine the prompt and the resulting dataset, but this prompt is solid enough for a proof-of-concept. It’s pretty accurate – I had an intern check a sampling of 100 job posts, which were 100% accurate. The LLM did not hallucinate any benefit strings and correctly categorized the benefits in all cases.

        I tried a few ways of visualizing this data and the heatmap is the most useful. Here’s the benefits data summarized by two-digit NOC. “NOC” stands for National Occupation Classification, and it’s a widely-used taxonomy for categorizing jobs. Workforce includes NOC data in the jobs dataset.

        A few highlights to mention here – 80 is the code for middle management occupations in natural resources extraction and agriculture, demonstrating the limitations of this particular data set. That’s one job at AMCO produce. Looks like they’ve got a nice benefits offering. 43, 44, and 45 are notable for how rare material benefits are. Those are “occupations in education, law and social, community and government services” – 43 is assisting occupations in education and legal and public protection, 44 includes home care providers, 45 is crossing guards. Each of these are represented by only a small number of jobs in the database. In the case of NOC 43, Workforce people tell me these are mostly firefighters in Windsor and they have excellent benefits, but there are no firefighter jobs on offer in this data set. That NOC is represented by two jobs for casual teaching assistants.

        Not as clean but perhaps more useful – here’s the benefits data summarized by Workforce sector code, which are collections of one to four two-digit NAICS codes. “NAICS” is the North American Industry Classification System.


        In this case, Workforce adds value to their jobs data by creating a precise sectoral profile for each employer. If a company does engineering for the construction industry, they get assigned a sector code for both construction and engineering. 

        I would caution against drawing hard conclusions from this data no matter what, because of the limited maturity of the proof-of-concept, but the main problem is that we’re working with a subset of a snapshot of jobs data from a particular point-in-time. It’s a subset because of data ingestion problems with one particular job board. I’m hoping to assemble a bigger collection of jobs in the future to get a clearer picture of this data.

        Anyway, that’s the work I did. I came away from it convinced that LLMs can be a valuable tool for data analysts in helping to extract meaning quickly and cheaply from unstructured data. I’m pretty excited to use this technique in my data projects moving forward.

        34 min
      • Canadian Halloween Data Special

        Our first ever Canadian open data Halloween special!

        The top 10 films on Netflix leading up to Halloween

        Halloween movies are a lot of fun, but are not a big part of the season for our family. I wanted to understand how that compares to other Canadians. There is only one streamer that provides at least some information on the viewing habits of their users, and that provider is Netflix. The top 10 films on Netflix for the week including Halloween and the week before are below. Films that are “Halloween themed” are highlighted in red.

        The data only goes back to 2021 because that’s when Netflix started providing it. Halloween themed movies make up between 1 and 4 of the top 10 during these Halloween weeks. It’s hard to compare the US and Canada directly though because other than Netflix originals, the films in the catalog are quite a bit different between the countries. Anyone that has taken a trip to the US and accessed Netflix there can immediately tell.

        No film made it to number one except for Night Teeth in 2021 where it took the first place in one of the weeks in 2021. It placed second in Canada where a Dutch film called “The Forgotten Battle” made it to number one. Other high ranking films were “The Curse of Bridge Hollow” and a documentary called “The Devil on Trial”. Interestingly, there are no repeats between years – which isn’t the case for Christmas where titles like “How the Grinch Stole Christmas” repeat.

        If you like this kind of data, watch out for Netflix’s “What we Watched” reports that come out every six months or so. These include statistics on all films and TV shows in the Netflix catalog, not just a top 10. Unfortunately it does not have country-by-country breakdowns which makes its usefulness to our podcast limited.

        Spooky Szn Maps

        Stumbling across Census Mapper’s Halloween-themed maps was a spooky szn treat. Census Mapper is an app developed by software and analytics company Mountain Math’s Data Scientist Jens von Bergmann along with Alejandro Cervantes and makes Canadian census data easily accessible in map format. For every census year since 1996, Census Mapper has produced 3 Halloween-themed maps: Haunted Houses, Trick-or-Treat Density, and Trick-or-Treat Onslaught. Let’s take a look at each of these (and one created by us of course!).

        The description on the 2021 Haunted Houses map reads, “Stats Canada calls them ‘dwellings not occupied by usual residents’, we say they are haunted and mapped the percentage of haunted dwellings. If usual residents don’t live there, who does? We think they are haunted by the ghosts of unaffordability and financial impossibilities, so usual residents have fled.” This is a hilariously witty way to describe dwellings occupied solely by foreign residents and/or by temporarily present persons, the latter of which do sound a lot like ghosts…

        While Census Mapper notes that “in Vancouver 7% of the housing stock is haunted… and other cities like Calgary, Toronto and Montreal have their fair share, with 5.4%, 7.4%, and 7.1% of housing stock haunted, respectively”, we of course looked at our home region of Windsor-Essex for fun. 

        At first glance, not overly haunted, BUT, take a look at that big brown spot – that’s Pelee Island, the southernmost piece of land in Canada, and also, evidently, one of the most haunted, with 67.4% of dwellings not occupied by usual residents! Time for an MMM Halloween field trip to Point Pelee.

        Census Mapper’s Trick-or-Treat Density and Trick-or-Treat Onslaught maps are equally fun. The Trick-or-Treat Density map shows the pockets in a neighbourhood that are likely to get the most trick-or-treater foot traffic, displaying the number of trick-or-treat aged children per km². Census Mapper considers children aged 3 through 14 as being of trick-or-treat age (highly debatable). The Onslaught map helps you stock up on the right amount of candy for trick-or-treaters in your neighbourhood (or more reasonably, helps you hide from the masses of children that might appear on your doorstep) by showing the number of children aged 3 through 14 per doorbell in each area. Census Mapper boasts over 150,000 Canadians use the Onslaught map annually for their Halloween planning!

        These maps seemed like a super fun way to visualize Census data, so feeling inspired, we made our own spooky map with Census Mapper – Scary Movie Freakouts. This map shows one-person households as percentage of household types, thereby displaying the areas in Canada down to the DA level with the percentage of folks most likely to have a major freakout after watching a scary movie then navigating the darkness of their home alone. Across Canada, Quebekers might be seeing the most scary movie freakouts with 35.1% of their households being one-person households. Thankfully for those in Nunavut, only 19.7% of households are one-person households. In our home Census Division of Essex, 28% of households are one-person households, including Frazier and Katie’s. If we disappear, you’ll know why!

        At the more granular DA level in our region, downtown Windsor has the highest cluster of one-person households. Luckily for Pelee Island with its many dwellings not occupied by usual residents, there are likely less major scary movie freakouts happening with 37.5% of households being one-person households. This map should compliment our Netflix Halloween movie roundup well for you scary movie fans who live alone and love the freak out feeling – you’re welcome.

        Public Opinion on Halloween

        I pulled several public opinion surveys about Halloween to see what polling firms and people are saying about the annual spooky season over the past few years. 

        First there is a large sample survey over 9,000 Canadians by “Caddle” that was done in partnership with with the Retail Council of Canada (RCC). Second is a survey from Leger from 2023 with 1,521 respondents and 2024 polling that came out just today with a similar poll size. Third is a survey from Ipsos for MARS with 2,000 respondents in 2024 that was released a few weeks ago. 

        Halloween surveying doesn’t seem to be a regular polling set, sometimes a question gets lumped in with other fall surveys but most of the major polling firms don’t seem to have a annual halloween survey which I kind of find surprising. It is likely they do, they just don’t publish results each year. I did find an older survey from Ipsos from 2020 that was surveying during COVID that said “COVID is the Grinch who stole Halloween” 

        As a baseline it is interesting as during COVID only about 1 in 5 Canadian households planned to give out candy. Atlantic Canada and the Prairies near 30% while Ontario was at 16% and Quebec with the lowest levels at 13%. In 2023, Leger found that this number was 48% of Canadians and in 2024 at 47%. In 2024 Ipsos found it to be 52%. You can certainly see the impact on the pandemic in the polling on whether or not candy was to be had over time.  

        A common question across all of the polls related to whether you were going to spend – more, them or less than compared to last year. Across the 3 polls Canadians planned to spend the same (66% and 71% in 2023, 68% in 2024). Leger went off the board and decided to ask about paranormal activities in their survey finding that: 

        As spending and affordability are top of mind issues in 2024, Leger published a full suite of data asking “How much will you spend on treats and candy this Halloween?

        The larger sample Caddle survey was more of an industry focused survey focusing on the spending patterns. Where people planned to buy their candy, how many purchases, when they expect to make them (weeks before Halloween), and how much they plan to spend and on what sort of items. 

        It is interesting that the data from this survey shows that Halloween is a “holiday” that is somewhat online resistant, with Canadians preferring going to stop and also willing to make an extra trip. 

        The Ipsos survey was interesting as they were able to separate out responses by generational age group. 

        The Ipsos survey in 2024 was funded by MARS so it isn’t surprising that the only candy choices by province/region were MARS brands 

        Overall the polling is interesting. The consistency in questions across polls and years does provide some interesting insights on Canadian spending patterns and preferences which is a little different that the usual political horse-race that is the usual conversations related to polls. . 

        Pumpkin Production in Canada

        You aren’t ready for spooky season without a jack-o’-lantern. The iconic carved lantern may once have been made from vegetables like rutabagas and turnips, but in Canadian Hallowe’en culture the jack-o’-lantern is universally a pumpkin. With that tenuous connection made, I set out to create a tool for querying and visualizing vegetable production in Canada, with the demo focused specifically on the mighty gourd of October, the pumpkin.

        Statistics Canada publishes data about vegetable production in Table 32-10-0365-01, “Area, production and farm gate value of marketed vegetables.” This table is sourced from the Fruits and Vegetables Survey of a random sampling of farms. Data is reported for several measures, grouped by vegetable commodity; 

        • Area planted
        • Area harvested
        • Total production
        • Marketed production
        • Value at the farm gate
        • That last one requires a bit of explanation – it’s the price of the commodity minus transportation and marketing costs. 

          These numbers are aggregated at the level of the province and reported annually. Territories aren’t included. Also excluded are institutional farms, community pastures, mushroom and potato farms, and greenhouses. 

          The table goes back to 1940, but data series for different commodities start at different times. We have pumpkin data from 2007. 

          Production volume across Canada has more than doubled since measurement began, and the inflation-adjusted price of a ton of pumpkins is up about 12% since 2007.

          When you break it down by province, it’s clear that Ontario is the undisputed pumpkin kingdom of Canada – in most years, Ontario produces more of the orange gourd than the rest of the country combined.

          In terms of value to farmers, pumpkin production was worth almost 40 million dollars in 2023. The chart below is not in constant dollars.

          You can find all the code, data, and visualizations for this segment on GitHub. I downloaded a CSV from the Statistics Canada table for Canada as a whole, and each province separately, gave them consistent file names and worked with my AI coding assistant to create a Python script to process the data files and generate a SQLite database, and another Python script that uses the matplotlib library to generate interesting charts from the data. The repository comes with a Dockerfile so you can work with the data in a container if you like that sort of thing. Finally, I threw in a few SQL queries to help explore the data.

          This repository will work for any commodity in the Fruit and Vegetable Survey if you swap out the data files. The project can pretty easily serve as the starting point for a more comprehensive agriculture data tool.

          There are a couple of technical points to talk about with this project. 

          I want to highlight the value and convenience of SQLite for this type of work. It’s fast, stores data efficiently and the database engine is pretty capable. The primary constraint to working with SQLite is that it’s really only good for one user at a time, but that’s no barrier for data analysis work like this. I did some exploring of the data using SQLite and Excel together. This is a really nice approach and I’m glad I tried it. You can paste into Excel directly from SQLite query output by using the Import Text Wizard from the Paste Special menu. That makes it easy to bring in external data for ad-hoc exploration, like this simple table applying a present-day inflation adjustment:

          The other point I want to mention is that as I continue to gain experience coding in tandem with LLMs, I learn more and more about the nuances of working with them. LLMs are GREAT at writing Python code but less capable of doing things with a smaller footprint in their training data. Try getting ChatGPT to build an Observable notebook if you want to see what I mean. Which isn’t to say you can’t pair with an LLM for more obscure technology, but you should be prepared to do a lot more oversight of your coding partner and to break your problem into smaller pieces for effective results. Also they’re pretty terrible at debugging, but so are most people so that’s not such a huge insight.

          40 min
        • Pipelines, PowerBI and Price Indexes

          This month on Mean, Median, and Moose, geeking out on tools.

          Creating a dashboard of LCBO Top 20’s with Observable Framework

          By John Haldeman

          For my tool, I picked the Observable Framework. This is an interesting cross between a static website generator and Observable’s JavaScript notebooks. Framework has a good getting started guide that will get you the feel for how this works.

          First you need some data to populate your data application. I chose an interesting project called LCBO stats. This project is an open source API which gets its data by scraping the LCBO website. You can then use the API to get rankings of LCBO products. From there you can look up their pricing history. Pretty cool!

          To extract the data from LCBO stats, I created a simple data loader which is just a little script that outputs a data file – in this case formatted as JSON. We’ll use a “FileAttachment” API later to have Observable Framework automatically execute this script and insert the data into the website. What’s interesting and unique about this is that the Observable Framework supports data loaders in a variety of languages. All you need is a simple program to get the data and dump it to standard out. My data loader takes two parameters, field (to specify what to rank by) and sort (to specify ascending or descending order), so that I can use a single one for all of my charts. I didn’t do any aggregations because the API doesn’t support it and I didn’t want to programmatically page through all the results and do the aggregations myself, because that would be a rude thing to do with a free public API (ie: use a bunch of resources to dump all the data available).

          I then displayed the data using a sort of “Markdown file on steroids” which the Observable Framework excels at. This includes a markdown header to specify the themeing and other options. A JavaScript block to retrieve the data and a block to display the data using Observable Plot. Finally you just insert the results from the JavaScript calls into some HTML.The ergonomics of this are amazing. No tagging HTML tags with IDs and telling the chart framework to dump the data there. The code and the presentation mix effortlessly. Traditionally that’s a bad idea, but the notebook-like nature of these definitions means that you can create complicated things and they’re still easy to understand.

          And here’s the end result: https://johnhaldeman.observablehq.cloud/lcbo-top-20s/

          Data Pipeline Demo

          By Doug Sartori

          Many people use Python for data transformation. Python libraries like PETL and pandas provide many of the tools needed for one-shot data analysis and transformation work. Data notebooks provide interactivity and make for a pretty compelling set of tools for data professionals.

          We’ve talked a lot about those tools on this show, which makes sense because they’re the basic tools you need to work the datasets we deal with every month. ETL is only one part of the picture, though, so with that in mind let’s zoom out and look at some of the options out there for data pipelines.

          Tech vendor IBM defines a data pipeline as a method in which raw data is ingested from data sources, transformed, then ported to a data store. The ETL tools and methods we’ve talked about are a key component of a data pipeline, but as data volume and complexity grows, you will increasingly feel a need for a tool to organize, orchestrate, and perform your ETL tasks.

          When Doug’s development team needed a job runner for a recent project, they landed on Sidekiq. Sidekiq is a solid job processing project with a lot of relevant features for creating a data pipeline. It runs tasks asynchronously, pulling jobs from a queue stored in a Redis data store. Notification and complex error handling are all supported well. Here’s a screenshot of the Sidekiq web UI.

          The project was a success and Sidekiq did everything they wanted it to. If you use Ruby, and your tasks are mostly independent of each other and don’t require modeling dependencies, Sidekiq is a good fit for organizations at many sizes. 

          A more robust feature set is found in the Mara project. It’s a loose set of Python libraries implementing many data warehousing features. Mara core modules provide access control, schema management, ETL tools, and pipelines. Along with a target data store, it’s potentially a complete solution for an organization’s data. Unfortunately, Mara doesn’t come with a lot of documentation. It’s pretty tough to build a working local copy of the Mara example system without reading Mara source code. Mara’s Github page has lots of activity and this code is clearly widely used, but adopting it requires a significant investment of technical capacity.

          For Doug’s use case, and maybe yours, Spotify’s Luigi project comes pretty close to the sweet spot. It’s a lot more robust and widely-used than Mara, and has features that Sidekiq lacks. Crucially, Luigi models workflows as directed acyclic graphs (DAGs), which allows for structured workflows with dependency management.

           Doug built a Luigi demo you can find in this GitHub repository. The demo uses Luigi to run a short workflow of dependent tasks and populate a local SQLite database from multiple sources.

          The scenario for the demo is a company that needs to compensate its employees for expenses incurred in travel across North America. There is a CSV containing a data set of random names and synthetic expenses in different North American currencies in the repository.

          The public data source for exchange rates in this demo is the Bank of Canada Valet API. It’s a really rich service with tons of useful data, including daily exchange rates.

          The demo is contained in a Docker container that starts up the Luigi Central Scheduler when the container is started. It’s configured to expose the Luigi service’s web interface, which provides a web interface for monitoring current and recent jobs.

          Luigi doesn’t itself schedule task execution, which may sound surprising since we’re talking about the Luigi Central Scheduler, but the service is concerned with scheduling the execution of tasks within workflows, not with firing off workflows. For that, the Luigi people recommend writing your own service or using an operating system scheduling facility like cron.

          The demo is configured with three tasks;

          • FetchRates loads exchange rate data
          • ImportExpenses loads expense data
          • GenerateReport generates a report
          • GenerateReport depends on ImportExpenses, which in turn depends on FetchRates. 

            In the container, running the main Python script adds a GenerateReport task to the Luigi scheduler, which then fires off the other two tasks in turn to satisfy the dependencies of GenerateReport. 

            Luigi’s scheduler determines whether a task is complete or not by checking whether its output exists. If the output is already in place, Luigi won’t start the task. In the case of the demo, a report CSV is the output of GenerateReport.

            There’s more detail in the demo README file. Check it out, we hope you find it useful.

            Creating a Dashboard in Power BI:

            By Rashmi Krishnamohan

            As a data analyst, I work with a variety of tools, but Power BI is my go-to. Why Power BI? It’s simple—it allows me to tell a story with data. Instead of staring at endless rows and columns, Power BI turns those numbers into interactive, engaging visuals. It’s like giving the data a personality, making it easier for users to explore, understand. Plus, it keeps things fun—because let’s face it, if you can make data interesting, you’ve won half the battle. Even if someone has never used Power BI before, they can quickly get the hang of it and dig into the data themselves.

            For the dashboard I created for this podcast, I downloaded the Data Science Salaries dataset from Kaggle—a fantastic resource when you want to get your hands on real-world data for practice or portfolio projects. The dataset includes fields like:

            • Work Year: The year of the salary data.
            • Experience Level: Junior, mid, and senior and executive level professionals.
            • Job Title: From Data Analysts to Data Scientists to Data Engineers (and many more..)
            • Salary in USD: Standardized salary values for easy comparison.
            • Company Location and Employee Residence: Where the jobs are based and where employees live.
            • Remote Ratio: How much of the role can be done remotely.
            • Company Size: Whether the company is small, medium, or large.
            • Anyone who works with data knows that it’s never clean straight out of the box. This dataset was no exception—it had missing values, duplicate entries, and some inconsistencies in how fields like experience levels and company locations were recorded. That’s where Power Query Editor comes in.

              Power Query is one of my favorite features in Power BI because it’s like having an inbuilt toolkit for cleaning and transforming data. I used it to remove duplicates, handle missing values, and normalize some of the abbreviations. What’s nice is that all of this happens within Power BI, so you don’t need a separate tool to clean your data. Once it’s transformed, you can jump right into building your visualizations.

              In my day-to-day work as a data analyst, Power Query is a lifesaver when dealing with large datasets that need wrangling before I can even start analyzing them. Plus, if I need to export the transformed data from Power BI, I can easily use DAX Studio to push it back into a CSV format.

               Building the Dashboard:

              After cleaning the data, I got to the fun part—building out the dashboard. One of the things I love about Power BI is how versatile it is in terms of design and functionality. Here’s a rundown of some key elements I included:

              1. Slicers: I added slicers to allow users to filter by year and experience level. This makes the dashboard dynamic—whenever someone selects a different year or level of experience, the entire dashboard updates to show relevant data. It’s an interactive way to explore trends and comparisons.
              2. KPI Cards: I used KPI cards to display high-level metrics like average salaries, highest and lowest salaries, and year-over-year comparisons. These cards give users a quick snapshot of the key data points they’re likely to care about.
              3. Salary Forecast: One of the cooler features of Power BI is the built-in forecasting function. For the salary data, I created a line chart that not only shows past salaries but also forecasts future trends based on historical data. This is a great feature when you want to project potential future outcomes, and it’s surprisingly easy to set up.
                1. Top 10 Jobs by Salary: I always like to give users something to rank, so I included a chart that highlights the top 10 data science job titles based on average salary. It’s interesting to see which roles are leading the pack and how salaries differ across positions.
                  1. Maps: (Why Maps? Because They’re Awesome)
                  2. I also included a map visualization in the dashboard. Now, I’ll admit—I’m a bit biased toward using maps in my dashboards because they look amazing and are incredibly useful for visualizing data geographically. In this case, I used a map to show salary distribution across different regions, using a color gradient to make it visually appealing and easy to understand at a glance.

                    I took it a step further by using bookmarks to create different map views. So, you can toggle between maps showing average salary, highest salary, and lowest salary—adding even more interactivity to the dashboard.

                    Feel free to explore the dashboard yourself here.

                    39 min
                  3. Murals, Messages, and Moods

                    Today on Mean, Median, and Moose you’ll be treated to data we collected or generated all by ourselves!

                    Asking ChatGPT to Create Some Sample Data

                    I have been playing around with ChatGPT and a few other online AI tools. I wondered how well it would replicate re-world data and preference at a population level. I wondered how closely it could replicate election results by randomly generating a series of poll responses. 

                    First thing, I did was I had it view election coverage from the Windsor Star, CBC Windsor, CTV Windsor for the 2021 Election. I copied the links to ChatGPT asked it to summarize each story and made sure that it noted how parties performed in Windsor West.

                    I asked Chat GTP to review a number of news stories that I found by googling Federal Government News Windsor 2024, against asking ChatGPT to summarize the stories and not the important policy issues highlighted in them.

                    The top political issues in Windsor, Ontario, include:

                    1. Healthcare: Improving access and funding for healthcare services, especially mental health care and home care services.
                    2. Housing: Increasing the supply of affordable housing, implementing rent controls, and addressing homelessness.
                    3. Economic Development: Boosting local industries, including the auto sector, and improving job opportunities.
                    4. Education: Enhancing funding for schools and addressing educational disparities.
                    5. Environmental Policies: Promoting renewable energy projects and sustainable development initiatives
                    6. Based on these economic priorities and past election results, I asked ChatGPT to create a fake dataset of 250 responses of a poll about the riding of Windsor West in Windsor Ontario Canada including respondents: age, gender, level of education and household income level. To give additional contest on the demographic figures I had ChatGPT summarize Windsor West Riding Profile for each of the respondent categories. 

                      Use past elections, survey data and local news as a basis to create the responses to the following questions:

                      Question 1: Did you vote in the last federal election?

                      Question 2: Which party did you vote for? 

                      • Liberal
                      • NDP
                      • Conversative
                      • Green Party 
                      • People Party of Canada
                      • Question 3: How likely are you to vote in the next election?

                        • Very Likely
                        • Likely
                        • Neutral
                        • Unlikely
                        • Very Unlikely
                        • Question 4: What is your top issue in the next election?

                          Although it offered a Python output that stumped me and I was lazy, so I asked for a CSV file. ChatGPT out 250 rows of data like this. 

                          Respondent ID,Age,Gender,Education Level,Household Income,Voted in Last Election,Party Voted For,Likely to Vote Next Election,Top Issue 

                          1,45,Male,Bachelor’s Degree,$50,000 – $75,000,Yes,Liberal,Very Likely,Economy 

                          2,34,Female,High School Diploma,$25,000 – $50,000,No,N/A,Likely,Healthcare 

                          3,29,Male,Master’s Degree,$75,000 – $100,000,Yes,NDP,Neutral,Environment 

                          4,54,Non-binary,Some College,<$25,000,Yes,Conservative,Very Likely,Education 5,62,Female,Associate Degree,>$100,000,No,N/A,Unlikely,Housing

                          You can view the data here.

                          Created Sample of 253ConservativeGreen PartyLiberalN/ANDPPeople’s Party of CanadaCount of Party Voted 353437111351

                          Based on ChatGPT review of history and estimate of the future we could expect 57% turnout in Windsor West in the next Election up from the 43% in 2021. 

                          Created Sample of 142 VotersConservativeGreen PartyLiberalNDPPeople’s Party of CanadaVote Percentage 24.6%23.9%26.1%24.6%<0.1%

                          Now this is an unweighted sample and digging through cross times, find significantly over samples higher levels of education and incomes for the Windsor West riding. The top issues also made me laugh for a bit. 

                          Row LabelsConservativeGreen PartyLiberalN/ANDPPeople’s Party of CanadaGrand TotalEconomy191112Education2141711760Environment27342660Healthcare5193161962Housing64542259Grand Total353437111351253

                          Conservatives don’t care about the economy, Greens don’t care about the Environment, only people who didn’t vote want housing. Seems an accurate representation of my riding (sarcasm). 

                          Collecting Personal Data (Katie)

                          I’ve been an avid user of Daylio, a mood tracking app, for over 5 years now. It started at a time when I felt a lot of anxiety and unrest, so I wanted to be more mindful of my moods. I’ve tracked my range of emotions every day for years, with a reminder on my phone popping up every 3 hours to choose my mood and activities I’ve been doing. Gradually, I noticed my moods transition from more negative to positive, and I mostly select “Good” as my mood these days. 

                          A little over a month ago, I was faced with a new reason to collect personal data like this – I had noticed my fatigue, a symptom of my multiple sclerosis, seemingly impacting me more and more. But how could I really be sure without any data to back this up? And if it was happening, were there any patterns I could recognize to help with it? So, I transitioned my Daylio mood tracking to fatigue tracking, changing the 5 mood levels to 5 fatigue levels instead: Exhausted, Fatigued, Neutral, Awake, and Energetic. I also added activities like “Meal”, “Coffee”, and “Snack” to see if there was a correlation between my fatigue and when I was eating or drinking coffee. 

                          Daylio offers a wealth of statistics and charts once you’ve been tracking for a week or more. While I only have fatigue data for the month of May so far, I took a look at the stats Daylio has to offer to see if my assumption that I’m feeling fatigued often is true, and to see if there was any correlation with the activities I had listed, which you can see below. 

                          In May, I made 244 entries in Daylio (approximately 8 entries per day, with an entry about every 2 hours from 7am to 9pm). My average fatigue rating was 3.1 (Neutral), with a total of 3 Energetic, 66 Awake, 139 Neutral, 34 Fatigued, and 2 Exhausted fatigue levels entered in the month. 

                          This told me that while I wasn’t doing terribly with my fatigue, I also wasn’t doing as well as I wanted. Ideally, I’d be “Awake” or “Energetic” over 50% of the time, not 28% of the time.  I entered “Neutral” so often that Daylio considered my fatigue quite stable though, scoring me an 86/100 in their stability chart. 

                          So, now I knew I was probably more fatigued than I’d like to be, and I took a look at the activities I had logged to see how these might be impacting me. According to Daylio’s “Most Influential Activities” chart, I was most awake after having a snack, shopping, having coffee (no surprise there!), when I was at work (much more surprise there), and after having a meal. Going for a walk, watching a movie, going to bed, traveling, and seeing my family were activities aligned with poorer fatigue levels. 

                          Now, here’s where I had to take some of this with a grain of salt. While I could see snacks, shopping, coffee, and meals helping me feel more awake, “Work” likely appeared on the list simply because it was my most logged activity at 98 entries, and I also almost always have a meal and coffee while at work. I could see this when clicking on “Work” as an activity in Daylio’s stat tracker and viewing “Related Activities” with “Meal” occurring 22 times on the same day as work and there being a 71% relation between the two, and similarly with “Coffee”, at 67%.  

                          Similarly and intuitively, I know “Going to bed” made the negative list since of course I’m more tired right before bed, “Movies” appeared negatively since I almost only watch a movie right before bed. In the same way, “Family” made the list as I always go to see my family Monday night right after work, when I’m naturally more tired. 

                          All in all, tracking my fatigue in this way helped me be more mindful of it, confirmed my average fatigue level, and helped me see that eating and a cup of coffee are ways I can boost my energy level, even if only temporarily, whereas exercising through a walk might not be as energy-boosting for me as it is for others. Tracking your own personal data can be simple with the wide variety of tracking apps out there now, and it’s a great way to get to know yourself better, increase your mindfulness, and tackle personal goals in a highly intentional and analytical way. 

                          Developing a Walking Tour

                          A few years ago, Doug’s company Parallel 42 Systems built a government-funded walking tour app. P42 requested and was granted the permission to commit the code written for this project to the commons. It’s called Pytheas, and since the code is free for anyone to use, they’ve been using it! Last year P42 collaborated with local LGBTQ+ activists on a tour of sites of historic significance to the queer community, and this year their gift to our local community is a curated cross-border art tour focused on murals.

                          The code is mostly written, so implementing a new Pytheas tour is all about data collection.

                          There are some data points needed before a mural can be included in the tour;

                          • Geographic coordinates
                          • A photograph of the mural
                          • The address of the nearest building
                          • Ideally, the data for a mural includes the following information;

                            • The name of the artist
                            • A brief biography of the artist
                            • A description of the work, its origin, and its meaning
                            • Murals are inherently ephemeral, and generally not well-documented. Data collection involved automated and manual steps. To get started, the team looked for existing data sets of murals in the two cities.

                              In Windsor, the Free for All Walls Festival in 2023 added dozens of murals to local streets. This project has an excellent website and map which provide a good starting point, and crucially for this project, artist information for each of the murals.

                              In Detroit, the Visit Detroit Mural Guide and the City’s Mural Map also provided some useful hints, though the City of Detroit’s map does not surface many of the key pieces of information we needed.

                              Starting with this seed data, Doug and the P42 team next determined the desired route for each tour based on background knowledge of each city and the objective of the tour, which is to promote cross-border tourism in the region.

                              P42 used social media to ask residents of each city their favourite murals, documented them along with the murals from the seed data that appear on our target route.

                              The next step was the development of a complete list of candidate murals. This work was performed by a pair of site surveys. The first survey was conducted via Google Maps Street View. Street View was used to find murals, and to understand the immediate environment around the murals. 

                              At this point, P42 had a list of about sixty candidate murals on either side of the border that fit the basic criteria. The spreadsheet of candidate mural locations was geocoded with Geoapify, which is a service we’ve talked about on this show before. That geocoded spreadsheet was uploaded to Google Maps for use by the photographers hired to walk the routes and get photos for the tour.

                              P42’s photographers used telemetry tools to capture a GPX document identifying the locations of the murals, and performed on-the-spot curation of the murals and the tour by being the first ones to walk it.

                              Using QGIS, P42 converted the GPX files to CSV, hand-modified them to reflect the structure and naming convention of the geoJSON files that power Pytheas, and finally converted them to geoJSON. Supplementary information from various sources was manually added to the final geoJSON document.

                              You can see the results at https://motownmurals.tours. If you’re not local to Windsor/Detroit, you’ll be too far away to take the tour, but you can browse the list of murals and check out the photos. 

                              The MMM group chat

                              Behind the scenes at Mean, Median, and Moose there’s a rollicking instant message group we use a little bit to coordinate the show, but mostly to post funny tweets, memes, and complaints about local politics. Given that we’ve been doing this for a few years, we wondered what kind of data the group chat itself could provide. To limit the scope a bit we analyzed all our messages from 2023 – all 29,272 of them. That’s an average of 80 messages a day. 20 messages per person per day….

                              Broken down by chat participants we see the most messages come from John, then Doug, Frazier and Katie in that order.

                              John is also the most prolific link sharer, but this time Frazier comes in at number 2:

                              What about the time of the message? We can see that May and June were the most popular months for posting:

                              December is the least popular – I think likely because this is when the Mean, Median, and Moosers might be busy with the holidays. Ironically some of our most popular episodes and posts are the Christmas specials.

                              Finally, we can’t do an analysis without a heatmap, so here’s where the Mean, Median, Moosers were most active according to day and hour. The weekdays below start on Sunday (numbered as 1).

                              44 min
                            • Cannabis, Cheech, and Chong

                              This month on Mean, Median, and Moose, we look at Canadian data on Cannabis.

                              Statistics Canada Infographics

                              I was a little behind this month so I dug into a couple of different data sets. First I looked up what infographics were available from Statistics Canada. From fall 2023 there was a graphic that compared booze to pot sales in Canada from 2021-2022. 

                              An archived graphic from 2021 showed how legalization scaled up cannabis 

                              These are the only two infographics from Statistics CAnada that I could find, there other government of Canada infographics related to cannabis – like this one for legal risk from Health Canada; impaired driving, boating, flying; 

                              Munchie Specials

                              I also did some digging into what local restaurants had specials for 4/20. Digging through 40 local restaurant facebook pages I found 3 who had some sort of 4/20 special. Two are local places, while two are chains who had these specials beyond our community. 

                              Cannabis Retail


                              In 2018, StatsCan launched the Cannabis Stats Hub, which has been discontinued and replaced with a page listing twelve data sets Statistics Canada publishes around Cannabis. Many of these, like the table of Cannabis consumer prices, have been discontinued since legalization.

                              Since legalization of Cannabis in 2018, Statistics Canada has reported an increasing number of data sets related to the production, distribution and sale of cannabis products in Canada.

                              Some of the currently-available data sets cover topics like prevalence of cannabis use, the value of cannabis produced in Canada, and household spending on cannabis. Besides these topics, there is plenty of detail on the retail sale of cannabis in Canada.

                              Retail sales are reported annually, broken down by province and type of cannabis. So far, there are two reporting years spanning 2021/22 and 2022/23. The data is provided in a number of different formats for download, including a format labeled “for database loading” which is terrific if you want to use SQL to query your data or drop it into Excel and manipulate it with pivot tables.

                              We did both. Here’s a screenshot of the Excel output showing sales by geography in each of the two periods. The data includes a summary value for all types of cannabis, so you have to be careful to filter out that value (or all other values) to get a valid result. 

                              You might notice that retail cannabis sales were zero in both periods for Northwest Territories and Nunavut. Great news for Nunavut tokers – there is now a single licensed retailer in Iqaluit. The Northwest Territories now has a whopping six stores licensed for cannabis retail. 

                              Here’s the code of a SQL query to generate a similar result, assuming you’ve loaded the data using the column names in the CSV file into a table called “CannabisSales”:

                              SELECT
                              geo,
                              Type_of_cannabis,
                              sum(CASE WHEN ref_date = '2021/2022' THEN value2 ELSE 0 END) AS '2021/2022',
                              sum(CASE WHEN ref_date = '2022/2023' THEN value2 ELSE 0 END) AS '2022/2023'
                              FROM CannabisSales
                              WHERE Type_of_cannabis <> 'Total cannabis products'
                              GROUP BY GEO,Type_of_cannabis
                              ORDER BY geo,Type_of_cannabis

                              Dried cannabis flower is by far the most popular type of cannabis purchased by consumers. It burns up a little over two-thirds of the dollars spent by consumers. Dried flower is followed by inhaled cannabis extracts, which are more commonly known as vaping products. They’re another 22% of sales or so. Edibles are notable among the “long tail” types of cannabis, eating up just under 5% of the market.

                              Statistics Canada also produces a data set breaking down the Net income of cannabis authorities, along with associated government revenues. In 2021/2022, governments realized about $1.2 billion in excise taxes, provincial sales taxes, GST and other revenues. In the following period that number was just under $1.5 billion, while retailers’ net income landed just under $2 billion – a healthy growth in sales and government revenue! This data is offered broken down by province and territory as well, if you want to drill down further geographically.

                              Cannabis sales are regulated provincially, so information about retail stores is provided at the provincial level. In Ontario, the AGCO’s data inventory contains a few data sets on the cannabis retail license lottery program and data on which municipalities opted in and which opted out of cannabis retail sales.

                              In Ontario, the AGO maintains a web page that provides a list of license applications by application status, which is also downloadable in CSV format. If you’re curious about the subset of retailer applicants currently in their public notice period, which offers local residents an opportunity to respond to the application, you can find that information on the AGCO site as well.

                              Using a handy geocoding tool, we converted this data into a list of geographic points and imported it into QGIS. It’s a bit sparse for a heat map, so we used QGIS point clustering functionality to show the number of cannabis retail outlets in a given community. You might notice an irregular line of single retail stations in Northern Ontario. That mostly follows the Trans-Canada Highway. There are only two retail stores in Ontario north of the highway.

                              Wackiest Tobacky Sold on Government Cannabis Websites

                              When cannabis was first legalized, it seemed most available products had a moderate amount of THC, with the highest percentage in the low 20s. As your average Canadian began their foray into legal cannabis, it could be safe to assume they didn’t immediately buy cannabis with the highest THC percentage they could find. Now, it seems the Canadian customer has been demanding wackier tobacky than in that first year or two, with higher and higher THC levels reaching into the upper 30 percent. To see just how wacky the tobacky gets in each province, we took a look at the government cannabis website in each province (as available!). Immediately on diving into this manual data scraping adventure, we learned there’s no government cannabis website in Alberta, Manitoba, Nunavut, Saskatchewan, or the Yukon, so those are excluded from this analysis. Additionally, while BC, PEI, and Quebec have government cannabis websites, there’s no option to sort by THC level (perhaps an attempt to keep their residents in their right minds!?), so those too have been excluded from this analysis on the account of it being a bit too time-consuming to scroll through pages of cannabis products in search of the highest THC level. This left New Brunswick, Newfoundland, Nova Scotia, Northwest Territories, and Ontario to duke it out for the wackiest tobacky sold on their websites. To keep this analysis doable and comparable, only cannabis sold as whole dried flower was considered. See the ranking in the very fancy table below!

                              ProvinceName, Origin, Brand, TypeTHCCBDPrice Per GramONMega Breath, ON
                              True Fire & Co Ltd.
                              Hybrid

                               33-39%
                               0-1%
                               $10.06/gNSBanana Mints, NS
                              Eastcann
                              Indica

                               30-38%
                               0-1%
                               $9.99/gAnimalZ, NS
                              Eastcann
                              Indica

                               30-38%
                               0-1%
                               $10.71/gNBOrganic Kiwi Banana Cabana, NB
                              Eco Growers Choice
                              Indica

                               33-37%
                               <1%
                               $13.14/gNTGelatti Kush, QC
                              West Island Culture
                              Sativa

                               30-37%
                               0-1%
                               $11.21/gNLDeath Star, SK
                              Bold Growth
                              Indica
                               30-36%
                               0-1%
                               $9.68/g

                              The aptly named Mega Breath, sold on Ontario’s OCS and produced in Ontario, takes the cake for having the highest % THC whole dried flower product available, ringing in at a whopping 33-39% THC. This True Fire product is described as “a powerful hybrid strain with a unique and potent blend of indica and sativa. The aroma of Mega Breath is pungent and earthy with hints of pine and slight sweetness that lingers in the air. Its taste is just as impressive with a smooth and complex flavour profile that combines notes of spice, diesel and citrus.” It comes with a price tag of $10.06/g, making it a pretty sweet deal. The OCS seemed to have the most information available of the government websites, with information like grow method, grow medium, and grow lighting listed, and craft cannabis stamps on products – there was far less information available on the remaining websites.

                              Banana Mints and AnimalZ are next in line, each sold on Nova Scotia’s NSLC website and produced in Nova Scotia. Each are indica Eastcann products with 30-38% THC. Banana Mints has a price tag of $9.99/g, and Animal Z is priced at $10.71/g. Though very similar in type, THC, and price, Banana Mints’ flavours are listed as “kiwi, banana, and earthy”, while AnimalZ has “nutty, berry, gas” flavours (someone please explain why gas is a desirable flavour!). The NSLC website was lacking descriptions for its products, but it did mark its weed with “Proudly Nova Scotian” if it was grown in the province, and you can check local store availability too!

                              Cannabis New Brunswick is third up, selling Organic Kiwi Banana Cabana, an indica with 33-37% THC produced in New Brunswick by Eco Growers Choice. It’s sold for $13.14/g, with the description reading, “The experience begins with intense aromas of kiwi complemented with hints of sweet citrus. The taste and flavour of the banana is as powerful as the immediate onset – which comes through with notes of melon, honey, and pineapple. Just like a ‘Copa Cabana’, who could ask for anything more… this strain packs a punch with a strong cerebral effect, followed by a hard-hitting trance of relaxation leaving you feeling heavy headed, calm and focused to take on the night with your friends on the beach.” Sounds nice (though expensive!)! New Brunswick also had by far the coolest website, with senses like “fruit” listed for products, local store availability, and even the ability to leave reviews for products! Good job, NB.

                              Fourth, we have Northwest Territories’ website, Releaf, which is actually just a privately owned chain with the only delivery available in the territory, so naturally, the government has designated it as their “official” cannabis website. Sadly, it seems it is close to shutting down as it can’t get enough of the right products for delivery, or so the website reads. Gelatti Kush, a sativa produced in Quebec by West Island Culture, has the highest THC on their website with 30-37% at a cost of $11.21/g. No description or flavours available for this one.

                              Last is CannabisNL, with the indica Death Star product from Saskatchewan’s Bold Growth available at a price of $9.68/g and coming with 30-36% THC. It has “diesel, earthy, and pine” flavours. No descriptions available on the CannabisNL website either.

                              There’s no real consistency in the type or brand of cannabis with highest THC offered on each website, though they do all have less than 1% CBD! Similar price points across the board too, minus Organic Kiwi Banana Cabana’s slightly higher price. It’s interesting to see the different approaches taken by each province to the cannabis rollout since legalization, and time will tell if we keep seeing higher and higher THC products released!

                              Cannabis related search results on the Libraries and Archives Canada catalog

                              Libraries and Archives Canada maintains the catalog for Canada’s National Archives and National Libraries. Anybody can search the catalog online. We decided to do some searches for Cannabis related topics and graph them by publishing data. First up, the obvious “Cannabis” search:

                              Interestingly the publications peak in the 1980’s rather than recent history, which you may expect given recent legalization. You may also think that 1780 is a little early for Canadian documents about Cannabis, and you’d be right. The lone record in the 1786 was a land petition for Lower Canada for a Mr. Benjamin Weed. Clearly archives Canada is smart enough to know “weed” is a synonym for “cannabis”, but not smart enough to know the proper name “Weed” is not – a much harder problem.

                              Let’s take a look at some Cannabis related cultural figures and how they’re represented in the archives:

                              Willie Nelson by far the most popular of the four. If you’re wondering what kind of Snoop Dogg related material our national archive organization stocks, the earliest published item for the artist is a CD insert for issue 179 of a weekly UK music industry publication called “The Tip Sheet”. “Snoop’s upside ya head” appears alongside a Celine Dion’s “All by myself” and a cover of Randy Bachman’s “You ain’t seen nothing yet” by an artist known as Loverman on the insert. Strangely, issue 179 appears to be the only issue of “The Tip Sheet” stocked by Library and Archives Canada.

                              41 min
                            • Coins, Colonists, and Cost Recovery

                              This edition of Mean, Median and Moose, we look at discontinued data sets!

                              Canada Year Book

                              Less of a data set, more of a product, is Statistic Canada’s annual Canada Year Book, in production from 2006 to 2012. Funny enough, this seems to have been a previously discontinued but revived product that was discontinued again, as it’s also available from 1867 to 1990! Billed as “the premier reference on the social and economic life of Canada and its citizens”, each year book is presented in almanac style with more than 500 pages of tables, charts, and analytical articles on every major area of Statistics Canada’s expertise. This was fondly used as a reference in many a high school research project.

                              The idea is, you can click on any of the chapters, ranging from “Business performance and ownership” to “Families, households, and housing” to “Prices and price indexes” and get a quick snapshot of Canadian life as it related to the topic that year. This made for a great quick-reference resource and one that could make for a fantastic starting point for exploring data and history, particularly for students.

                              It is interesting to see how the year books changed over the years, even from just 2006 to 2012. For example, 2006’s “Education” chapter provided much more commentary than 2012’s Education chapter, commenting generally on schools’ serving special needs students, the denominational system and its abandonment and uptake by province, immigrants lifting the education level in Canada, and the financing of education. In contrast, 2012’s Education chapter simply lists statistics on student enrolment, graduation, tuition costs, and adult training, with little narrative behind these statistics. Particularly eye-catching given the current tuition situation in the 2006 chapter is: “Undergraduate tuition fees have almost tripled since the early 1990s. In 2004/2005, university tuition fees averaged $4,172, compared with $1,464 in 1990/1991.” Seems university tuition has been on a steady rise even longer than our memories might allow, reaching an average of $5,366 in 2012 according to that year’s Year Book.

                              Another interesting event to see reflected in the series? The 2008 Recession, featured prominently in the 2009 “Business performance and ownership” chapter, the “Economic accounts” chapter, the “Income, pensions, spending and wealth” chapter, and many others. Again, this earlier Year Book takes more of a commentary approach, with the “Economic accounts” chapter introduction reading, “Until 2008, Canada had gone a record 16 years since its last economic downturn and had been riding a seven-year boom in commodity prices. But the economy in 2008 was unlike any in recent memory. For many younger workers and investors, 2008 was their first experience with a recession.” While younger Canadians were experiencing their first economic downtown in 2009, they were also disproportionately “browsing, blogging, chatting, and downloading” compared to other Canadians as the “Information and communications technology” chapter details (and at the high speed of 5-9 mbps!). If you can believe it, StatsCan mentions 500 internet service providers operating in Canada at the time!

                              These are such digestible and interesting accounts of Canadian history, and it’s a shame this time capsule series has been discontinued!

                              Canadian Coin Issuance

                              The Royal Canadian Mint has a website that contains all the numbers for the mintages for various Canadian coins. You can go, learn about the histories of each coin type and then see how many coins were issued every year. Unfortunately none of this data is in an easy to consume format, but never fear, Mean, Median, and Moose are here with their document.querySelectAll() super powers. You can find an easy to consume CSV file and some graphs with the numbers on the Observable notebook here.

                              So, what does this have to do with discontinued data sets? Well, this is a data set about something that was discontinued – the penny – but the data set is alive and well. See what we did there? The last penny was minted in 2012, but before that it constituted the bulk of Canadian minting in terms of number of coins:

                              The youngsters or new immigrants reading this might also be interested to know that the toonie is a modern invention, making its first appearance in 1996:

                              You might also be interested to know that the loonie was not substantially minted until 1987. Doug still yearns for the old school $1 bill but will have to settle for crossing the border to see one in the modern age.

                              Speaking of that initial mintage of the toonie in 1996, you can see that it’s a big one. So large that, in terms of face value, that year dwarfs all others:

                              $2 being worth 200 pennies makes that bar the biggest by far. If you look at the coin and value numbers in general though you can see there’s been a big reduction since about 2013 even though the economy is larger. As you’ve probably guessed, there’s a lot less coins being issued since we do fewer and fewer transactions with hard currency.

                              Reporting and Trends on Data Gaps 

                              There are literally dozens of discontinued Statistics Canada datasets that I could talk about:  ending of annualized tracking of marriage and divorce rates; to shifts in Census methodologies;  to more eclectic data sets like Salaries and salary scales of full-time teaching staff at Canadian universities ending. To reliability issues emerging in critical economic surveys: here in 2016 and here in 2023 as fewer Canadians complete the surveys. 

                              Statistics Canada actually has a page where they answer some common questions on “Does Statistics Canada Collect this information”. Some of these data items are certainly “nice to have” – dog and cat pet data or the proportion of the population that is vegetarian or vegan. Others seem more critical – stats on abortion rates, homelessness, classroom sizes in schools. 

                              Reporting was done by the Globe and Mail in 2019. As part of a broader series on comparing Canada’s data landscape to other countries, they created an interactive tool for readers to ask their data questions, they identified 30 gaps in 2019 – ranging from how many people live in Nursing homes in Canada (Statistics Canada doesn’t know and there is no centralized count) to Eviction rates (they looked at an innovative student from Princeton as a potential pathway forward. 

                              Statistics Canada 2022-23 Department Results report breakdowns the activities and costs the Statistics Canada undertook in a particular year. To a degree this kind of replaces some of the year books that Katie was talking about as the documents breakdown by different topic areas a summary of reports and studies that were accomplished as well as a few meta-narratives. With over 30+ pages dedicated covering the core services this annual document is not lean on what they covered. 

                              More interestingly they breakdown their spending and the led me to noticing acknowledgements like this recent study by Statistics Canada had this under its title.

                              The study on Food Insecurity in Canada was released in November of 2023 uses data from the 2021 Canadian Income Survey to gain a better understanding of food insecurity, with a focus on families both below and above the poverty line and across income quintiles. The study also uses data from the 2019 Survey of Financial Security to examine the net worth of families who are more likely to be food insecure. Now sponsorship and cost recovery isn’t completely new for Statistics Canada. 

                              That being said, the question you have to ask is when over $500 million per year in revenue why are there still data gaps? 

                              Census of New France, 1665-1666

                              Although there are still censuses in Canada, we count this one as a discontinued data set because the French colony of Canada as a component of New France ceased to exist in 1763. The Borealis data repository is an academic data repository in Canada. The repository hosts a collection of pre-Confederation census data that is available to the public. Though we didn’t make any maps this time around, there is historical GIS data available in this repository, which creates the potential for some really interesting data projects!

                              The first census in North America happened under the administration of Jean Talon, who was the Intendant of New France at the time. This was a newly-created position responsible for the entire civil administration of the colony, including statistical data. By instituting the first census and in some cases personally conducting the census door-to-door, Talon earns the title of the first official statistician in Canada.

                              We made an Observable notebook visualizing some key findings from the census, particularly the demographic information that led Talon to institute a program importing young French women to pair with unmarried colonists. This policy is deeply connected with another element that will strike a modern viewer of this census data: there were certainly many thousands of residents of the territory called New France not counted in this census because they were indigenous people, and the “shortage” of women in New France was a consequence of French policy discouraging marriage between French settlers and the indigenous population.

                              We also included a heat map and data viewer for the fascinating data on professions and trades in New France. The list of occupations is interesting in itself – shoemakers are distinguished from wooden shoemakers, an important enough difference in 1665 for there to be two categories. The biggest categories are “Carpenter”, with 35 people following this trade in the colony and “Servants,” with dozens of servants in every region. Another notable category is “Gentlemen of Leisure,” of which 15 resided in Quebec at this time.

                              37 min
                            • Cilantro, Crust, and Crullers

                              This month on Mean, Median, and Moose we look at data related to fast food in Canada. Aside from the regular discussion about data, make sure you listen to the interview at the end with Saskatchewan open data extraordinaire Andy Dyck.

                              Statistics Canada on Fast Food 

                              I was a little late planning this month and so I went back to my safety blanket of Statistics Canada. I had some hope about this topic as in June 2023 they produced a report called “Is Canada Becoming A Fast Food Nation. Annual sales at restaurants reached $7.7 billion per month as of April. Statistics Canada does not explicitly state “Fast Food” rather compare full service restaurants, limited service restaurants, drinking places and specialty food services. Generally “fast food” falls under the limited service restaurant.

                              What we see in the data is that Canadians are divided over their restaurant preference.

                              The impact of the pandemic is clear on full service restaurants’ recipes. While fast food stayed pretty stable, the decline in sit down eating is clear. By April of 2022, spending had “returned to normal” and saw the two restaurants matching each other in sales. Another interesting item is that in non COVID years, there seems to be some seasonality with higher spending in summer months than winter months. 

                              Another way to sort of look at “fast food” data is through average consumer spending. The Annual Household Spending Survey (not annual anymore) tracks how average households spend on a variety of necessities, products and services. This includes two classifications of Restaurants: Meals as well as Snacks and Beverages. It is likely there are some fast food establishments captured in the Restaurant Meals category while I suspect the snack and beverage category is fast food by definition. 

                              When we break out the provinces, we do see some interesting variations.

                              NFLD, Sask and BC all saw increased average spending on meals at Restaurants (likely from) Restaurants despite COVID. It is curious that only Ontario and Manitoba have been on a consistent down trend bucking the major of countries that saw their peak in 2019 before retreating in 2021. 

                              Finally there is a very obscure dataset on E-commerce sales as a percentage of total sales (and value) available on Statistics Canada Website where you can clearly see the disruption of COVID-19 on the Food/Restaurant space. 

                              Tim Hortons Locations

                              For this month’s data set we did some web scraping of the Tim Hortons locations website to get a data set of all the Tim Hortons locations. We found 3,943 locations. Here’s the top 40 cities ranked by the number of Tim Hortons locations each has:

                              If you were expecting a simple population graph, you were just about right. It’s fun to look at the anomalies though. Mississauga has as many Tim Hortons as Edmonton even though it is 70% the size of Edmonton in terms of population and 86% denser. Scarborough, a Toronto district is listed as a separate entity, skewing its numbers lower than they should be even though it is still ranked as first. Dartmouth, NS has about as many Tim Hortons in Guelph and Cambridge Ontario even though the city is about half the size. Sudbury has as many Tim Hortons as Regina for a city three quarters as large. BC, Canada’s third largest province has just three cities breaking the top 40.

                              We then decided to convert the addresses to latitudes and longitudes in order to generate a map showing a heatmap of Tim Horton’s locations. Here’s the result:

                              Toronto is indeed the center of the Tim Horton’s universe. Here’s a closeup of Southern Ontario:

                              And just for fun, Alberta:

                              This got us thinking. What are the two closest Tim Hortons to each other? According to the Tim Hortons Website address list, that would be 1515 Main St East, Milton ON L9T 0W2 and 3025 James Snow Pkwy N, Milton, ON L9T 7S3 which according to Google’s geocoding APIs are 1.2 meters away from each other, but different addresses. Turns out the Tim Hortons is part of an Esso station on the corner that just happens to occupy two addresses. If you remove like that by finding the closest two Tim Hortons more than 100 meters away from each other, you find that the nearest Tim Hortons are across the street from each other on King Street West in downtown Toronto. Interestingly Google Maps doesn’t seem to know about the one at 150 King Street, only the one at 145 King Street – perhaps it’s a part of an corporate office:

                              The Tim Horton’s website says it’s there though if you look hard enough!

                              Vegetarian Options at the Top 5 Largest Fast Food Chains

                              As a vegetarian, it can be harder than you’d think to find a tasty option at one of Canada’s top fast food chains. While some chains have strived to provide more veg-friendly options in recent years, others have decided to steer clear of the vegetarian market entirely. According to ScrapeHero, the top 5 fast food chains by number of locations in Canada are Tim Hortons, Subway, Starbucks, McDonald’s, and A&W. To see which of these chains is the most veg-friendly, we took a look at their menu to see how many of their main lunch or dinner options (at least 380 calories+) were vegetarian-friendly and which chain gives you the best bang for your vegetarian buck. Of course, menus can differ across Canada, so we looked at the mains (with no customization applied) on each chain’s menu at their location closest to the center of Toronto (which is the center of Canada of course).

                              McDonald’s rings in with the highest number of main options on their menu at 49 options, but incredibly, has not one vegetarian option on their extensive menu. Subway and Tim Hortons have a very similar number of mains on their menu at 34 and 28 respectively, and each has 5 vegetarian options on their menu. While A&W has 26 mains, it has only 1 vegetarian option. Starbucks, as more of a coffee shop than food service chain, has a very small menu with only 6 mains listed, but of these 6 mains, 3 are vegetarian! This means proportionally, Starbucks by far wins out as most veg-friendly with 50% veg options, followed by Tim Hortons at 18%, Subway at 15%, A&W at 4%, and McDonald’s at a big fat 0%.

                              Chain# of Mains# of Veg Options% of Veg MainsMcDonald’s4900%Subway34515%Tim Hortons28518%A&W2614%Starbucks6350%

                              Now, which of these chains might offer the best “calories for dollar” value item if you’re on the hunt for a vegetarian lunch or dinner? Tim Hortons’ grilled cheese melt comes out on top at 500 calories for $5.79, which gets you 86 calories per dollar paid. It takes second and third places too with the cilantro lime veggie loaded wrap at 530 calories for $6.79, getting you 78 calories per dollar, followed by the habanero veggie loaded wrap at 500 calories for $6.79, getting you 74 calories per dollar. Coming in last place is Subway’s 6” veggie patty sub at 390 calories for $7.99, getting you 49 calories per dollar paid (who wants a veggie patty anyway?).

                              ChainMainCaloriesPriceCalories/DollarSubwayGreen Goddess Veggie Wrap750$10.4971Green Goddess Veggie Bowl670$10.49646” Green Goddess Sub490$7.79636” Mozzarella Bella Sub530$8.49626” Veggie Patty Sub390$7.9949Tim HortonsGrilled Cheese Melt500$5.7986Cilantro Lime Veggie Loaded Wrap530$6.7978Habanero Veggie Loaded Wrap500$6.7974Cilantro Lime Veggie Loaded Bowl560$7.9970Habanero Veggie Loaded Bowl530$7.9966A&WBeyond Meat Burger500$8.2960StarbucksCrispy Grilled Cheese450$6.4570Apples, PB, & Trail Mix Snack Box390$6.2562Tomato and Mozzarella Sandwich380$7.4551

                              Seems like Tim Hortons takes the cake for a good amount of vegetarian choices while also offering the highest value. Subway and Starbucks are also great options for veggies, and A&W will work in a pinch. Vegetarians should avoid McDonald’s at all costs.

                              Pizza Maps

                              In our neck of the woods, pizza is always a hot topic. Windsor, Ontario prides itself on its local pizza style just as neighbouring Detroit does. The excellent mapmaking community at DETROITography made a city map a few years back identifying the “territory” of chain and local pizza places. Their method was intriguing and easy to apply – they identified pizza place locations using health department restaurant inspection records. It’s an approach that works anywhere the local health unit publishes this sort of data, so Doug used the same approach in his local community to map out our own pizza places.

                              The Windsor-Essex County Health Unit publishes food safety inspection records on its website. Unfortunately, this data is not available in downloadable form, and the structure of the website makes automated screen scraping impossible. Fortunately, for a small data set manual scraping is a viable technique. Searching for “Pizza” in the facility name produces five pages of results. Copying-and-pasting each data table page into Excel, then manually searching for known pizza-selling establishments without “Pizza” in the name returns a reasonably complete data set of this sort of restaurant in the region.

                              The resulting spreadsheet contains restaurant locations defined by a street address, which we’ll need to convert into map coordinates. Doug used a tool called Geoapify, which provides good geocoding results by combining data from multiple sources. Results include a confidence level and source identification. 

                              Armed with this spreadsheet, Doug used QGIS to map the data, using Statistics Canada geographic boundaries to supply a base map and polygons for analysis. Joining these two layers by processing them using the QGIS built-in “Join attributes by nearest” tool. This tool connects a point layer, like our list of pizza places, with a polygon layer and adds attributes to the merged layer to include data about the nearest point to each polygon, including distance from the identified point. 

                              The Windsor Pizza Map is pretty colorful and shows the variety of pizza experiences available to residents of Windsor. The distribution of pizza places follows transportation arteries and population density. A majority of pizza territory in the city is marked “other” here, representing the 41 pizza places in the city that have one or two locations. Roughly thirty percent of the city’s pizza offerings are locally-owned and operated “mom and pop” shops, a number that increases significantly when you recognize there are also many highly successful local chains like Antonino’s, Armando’s, Naples, and so on.  

                              Windsor likes to think of itself as the pizza capital of Canada, and this map helps make a decent case. If you’re curious what your city’s pizza (or shawarma, or coffee …) map might look like, you might be interested in following along with Doug’s tutorial on YouTube.

                              1 hr 1 min

                            About Mean, Median, and Moose

                            From the publisher's feed

                            A podcast about Canada, data, and visualization