People Analytics Deconstructed

People Analytics Deconstructed

By Millan ChicagoBusiness
Download on the App Store

People Analytics Deconstructed episodes

  • Data Cleaning, Part 2

    In this episode, co-hosts Jennifer Miller and Ron Landis continue their discussion on the importance of data cleaning and management. They review three of five aspects of data cleaning that are critical to checking prior to the analytic phase. In this episode, they discuss how to check for linearity and normality, outliers, and multicollinearity.  

     In this episode, we had conversations around these questions:  

    • How do you check for linearity and normality in a data set?  
    • Why is normality important to check for in a data set?  
    • What are outliers? This includes both univariate and multivariate outliers. 
    • How do you identify outliers in your data?  
    • What are some ways to handle outliers?  
    • What is multicollinearity?  
    • Why is multicollinearity important to check and consider during the data analytic process?  

    Key Takeaways:  

    • We should always consider the distribution of a variable with respect to our expectations regarding the distribution. If the distribution is inconsistent with what we expect, we should devote time and energy toward understanding why. In cases where our ultimate analyses require assumptions of normality, we need to ensure that our data are consistent with that assumption. We may elect to transform our data on the basis of these analyses, but should always be able to explain why we have done so. 
    • Outliers are cases inconsistent from other cases. In the univariate case, these are scores that are either extremely high or low. In the multivariate situation, we inspect the "profile" of scores across measured variables to assess the degree to which the case is consistent with others. Once cases are identified as outliers, then determining what to do with them is important. Our discussion focused on some common ways of dealing with outliers. 
    • Multicollinearity exists when two or more of the predictors are moderately or highly correlated. This is typically of concern when conducting analyses using the multiple regression framework. Specifically, we need to assess the degree to which predictor variables are overly redundant (highly correlated) prior to including them in our models. The variance inflation factor (VIF) or tolerance are commonly used to assess multicollinearity.  


    33 min
  • Data Cleaning: Part 1

    In this episode, co-hosts Jennifer Miller and Ron Landis discuss the importance of data cleaning and management. They identify five aspects of data cleaning that are critical to checking prior to the analytic phase. In some cases, data management is often embedded in the data encoding and storage process (I.e., certain rules are in place to ensure that data fields can only handle one type of data such as a date). In this episode, they discuss how to check for data accuracy and what to do with missing data.  

    In this episode, we had conversations around these questions:  

    • What is data cleaning?  
    • Why is data cleaning important prior to the analytic phase?  
    • What are the five steps of data cleaning?  
    • How do you check for data accuracy in a data set?  
    • What does it mean to have missing data?  
    • What are some of the ways that you can evaluate your missing data?  
    • What do you do with missing data?  

    Key Takeaways:  

    • Data management is imperative to the data analytic process. Without a strong focus on the management process, the analyses and subsequent interpretation and use may be misleading and incorrect. While this topic may seem boring or perhaps intuitive, it is necessary to have a plan for data cleaning.  
    • There are five broad aspects of data cleaning. Some of this depends on the data and focal question but in general, some or all of these steps should be considered when conducting analytics. As noted above and also in the episode, some of these steps may be more automatic due to the platform and storage restrictions. The five steps include checking for data accuracy, missing data, linearity and normality, outliers, and multicollinearity.  
    • Data accuracy refers to whether the data are accurate and conform to the fields in which they are included.  
    • A missing data analysis checks for missing values. Depending on the type and kind of data, there are various procedures for handling these missing values. 
    34 min
  • Applying Multiple Regression to Test for Moderation

    In another technically focused episode, co-hosts Jennifer Miller and Ron Landis discuss how to use multiple linear regression to test models involving moderation (or interaction). In episode 18, we discussed multiple linear regression in which we used multiple variables to predict the outcome or criterion variable. But what happens if you have a situation in which the relation between the predictor and outcome variable is actually dependent upon (or is conditional upon) the level of a third variable? In this episode, we deconstruct moderation and some applications of moderation.  

     In this episode, we had conversations around these questions:  

    • What is moderation/interaction?  
    • Why might we want to use multiple linear regression (as opposed to analysis of variance, ANOVA) to test for moderation?  
    • What are some applications of moderation in People Analytics?  
    • What's the best way to communicate moderation results?  
    • What are some of the concerns when presenting visualizations depicting moderation?  

    Key Takeaways:  

    • Moderation or interaction involves evaluating with the relation between a predictor and outcome variable is dependent (or conditional) on the level of a third variable. For example, we might be interested in whether employee engagement predicts jobs performance. In this case, we have a simple linear regression. If we add a third variable, such as working environment (I.e., remote or hybrid), we can now ask whether the relation between engagement and job performance is the same across different working environments.  
    • Moderation and interaction can be used interchangeably. One can use regression based approaches or ANOVA to test for the presence of interactions, though regression allows for the use of continuous predictor variables.  
    • Moderation is an application of multiple linear regression. In multiple linear regression, the effects are additive meaning that each variable contributes additively to explaining the outcome variable. In moderation, the effects are multiplicative in that a product term needs to be included in the model to examine whether the variance explained in the outcome variable is over and above the when each variable is added independently to the model.  

    Related Links  

    • Millan Chicago 

     

    33 min
  • What is Machine Learning?

    In this episode, co-hosts Jennifer Miller and Ron Landis discuss the emerging field of artificial intelligence (AI). In particular, they discuss machine learning and two broad categories of algorithms, unsupervised and supervised learning.  

    In this podcast episode, we had conversations around these machine learning questions:  

    • What is artificial intelligence?  
    • What is machine learning?  
    • What are some applications of machine learning in People Analytics?  
    • What is the difference between supervised and unsupervised learning?  

    4 Key Takeaways on Machine Learning

    • AI is the field of computers simulating human capabilities to process data. Several examples of AI exist in our everyday environment including products like Alexa and Siri and other processes like financial detection fraud, purchasing recommendations, and driverless cars.  
    • Machine learning helps automate the analytic process. Supervised learning is an approach that predicts or classifies outcomes via "labeled" datasets. In this approach, the user has to determine the outcome and inputs that are used by the algorithm. Regression is a common type of supervised learning.  
    • Unsupervised learning is an approach that uncovers hidden patterns in the data utilizing "unlabeled" datasets. In this approach, the user does not contribute to the initial model building process. Cluster analysis is one example of an unsupervised learning technique.
    • Ron and Jennifer discuss how machine learning can be used in the context of People Analytics.  
    31 min
  • What is Multiple Linear Regression?

    Earlier in this season, we discussed a commonly used technique called simple linear regression. In this technique, we used one variable to predict an outcome. But, let's face it – life is a little bit more complex than just having one predictor and many times, organizations have lots of data that can be used to predict an outcome. In another technically focused episode, co-hosts Ron Landis and Jennifer Miller deconstruct multiple linear regression. They focus on using multiple predictors to predict a single criterion variable.   

     In this episode, we had conversations around the following multiple linear questions:  

    • What is multiple linear regression?  
    • What are some applications of multiple linear regression?  
    • What are some of the ways in which models can be built using multiple linear regression?  
    • What is mediation and moderation?  

    2 Key Takeaways on Multiple Linear Regression

    • Multiple linear regression uses multiple variables to predict an outcome (I.e., criterion) variable. The ultimate goal is to explain the variation in the criterion variable. One aspect to consider in this analysis is the relation between variables; that is, to what degree do the predictor variables correlate and how does that relation predict the outcome variable. Depending on the relation between predictors, either partial or full redundancy might be present. 
    • Ron and Jennifer discussed three questions that can be asked using multiple linear regression. First, you can assess the effects of particular predictors while controlling for others. Second, you can compare different sets of variables to find the most efficient model. Third, you can test for moderation and mediation.   

     Related Links  

    • Millan Chicago 
    • What is Linear Regression? 

     

     

    32 min
  • When Simple Statistics Have Big Impact

    In this episode, Ron Landis and Jennifer Miller deconstruct the importance of utilizing descriptive statistics as the foundation of starting the data analytic process. As many advanced statistical techniques are built on descriptives such as the mean and standard deviation, it is imperative to understand the characteristics of the data set being analyzed. 

     In this episode, they have conservations around the following questions: 

    • What are the various ways in which central tendency is used to understand the nature of a data set? 
    • What are the advantages and disadvantages of using different measures of central tendency? 
    • What are the different measures of dispersion? 
    • What are some contexts in which certain measures of dispersion should be used? 

    Links 

    • Exercise
    • Exercise Solution
    32 min
  • Questions to Consider when Designing Visualizations

    In this episode, Ron Landis and Jennifer Miller deconstruct the key characteristics to consider when developing visualizations. In working with data, many are faced with decisions about how to communicate results. Given that one of the primary functions of analytics is to inform various stakeholders of the results, visualizations and other representations of data often play an important role in communicating findings.  

     In this episode, we had conversations around these questions:  

    • What are some of the best ways to design visualizations?  
    • What are the best practices when designing visualizations?  
    • What are some ways in which visualizations can be improved?  

     Send us your questions! We're interested in answering people analytic questions! Let us know what challenges or opportunities you're currently working on. You can either send us a description or record a short audio file and send them to info[at]millanchicago.com. We will answer questions in future podcast episodes.  

    35 min
  • Analytics in Practice: How to Utilize Data Analytics for Performance Assessment

    In the first "Analytics in Practice" episode, co-hosts Ron Landis and Jennifer Miller deconstruct how to utilize the data analytic process for performance appraisal. Given the widespread and varied use of performance assessments in organizations, there are numerous opportunities to reap the benefits of applying data analytic thinking to the process.  

    In this episode, we had conversations around these questions:  

    • What are some of the decisions to consider for each step of the data analytic process in the context of performance assessment ?  
    • What are some of the ways in which performance assessment can be improved by thinking about through the lens of data analysis? 
    • Is there information collected during performance appraisal that could be used in ways to learn more about employee performance? 
    • Data analytic thinking takes place well before the actual analysis. We talked about how many of the choices we make when conducting performance assessments impact the data we ultimately can use. 

    Key Takeaways:  

    • Performance appraisal involves numerous choices that can be informed by taking a data analytic perspective to the process. 
    • The process of more fully using data analytics within performance assessment is something that HR departments can address in a building fashion. That is, we can steadily build performance assessment on a data analytic foundation based on specific organizational goals and objectives. 

     Send us your questions! We're interested in discussing real challenges and opportunities in the people analytic space! Let us know what people analytic challenges or opportunities you're currently working on. You can either send us a description or record a short audio file and send them to info[at]millanchicago.com. We will answer questions in future podcast episodes.  

     Related Links  

    • Millan Chicago 
    • Landis article - Selecting response anchors with equal intervals for summated rating scales
    35 min
  • What are Predictive Models?

    In this episode, co-hosts Ron Landis and Jennifer Miller deconstruct building predictive models and specifically, utilizing forecasting in organizational context. 

    In this episode, we had conversations around these questions:  

    • What are different types of data analytics?  
    • What are some of the decisions to consider when building predictive models?  
    • What are some contexts in which predictive models can be used in organizations?  
    • What are some of the data analytic requirements needed to utilize forecasting in organizational contexts?   
    • What are some clear steps that HR professionals can take to use predictive models?  

    Key Takeaways:  

    • In general, we can think about three broad categories of data analytics: descriptive, inferential, and predictive.  
    • Ron and Jennifer provide a framework of how to build predictive models. First, all the relevant variables and relations among those variables need to be in the model. Second, the model needs to have data divided into a training set and test set to determine how well the model predicts the data. Third, they discuss how the model can be used in organizational contexts.  

     

    Related Links  

    • Millan Chicago 
    32 min
  • What is Natural Language Processing?

    In this episode, co-hosts Ron Landis and Jennifer Miller deconstruct natural language processing (NLP), a technique used to drive insights from text based information. They focus on how natural language processing can uncover information from different types of text such as performance management reviews, employee engagement responses, pulse survey responses, and job descriptions.  

    In this episode, we had conversations around these questions:  

    • What is natural language processing?  
    • How can natural language processing be used in HR?  
    • What are some of the data analytic requirements needed to use natural language processing?  
    • What are some clear steps that HR professionals can take to use natural language processing?  

    Key Takeaways:  

    • Natural language processing utilizes machine learning algorithms to interpret and process text.  
    • Ron and Jennifer provide an in-depth example of how performance feedback in the form of text could be used in conjunction with quantitative ratings. They discuss the example in the context of the data analytic process. First, what problem are you trying to solve? Second, what kind of data do you have to answer the question? Third, they discuss some of the NLP techniques. Finally, they provide recommendations on interpretation and communication to other key stakeholders.  
    • At the end of the episode, Jennifer and Ron recommend steps for folks just starting out in this space all the way to the more advanced HR professional.  

     

    Related Links  

    • Millan Chicago 
    • Basics of Text Analysis for HR 
    • What is Natural Language Processing, and How is it Used in Workforce Analytics? 

     

     

    35 min

About People Analytics Deconstructed

From the publisher's feed

Are you responsible for understanding an employees’ experience? Have you tried to incorporate people analytics in your organization but have struggled? Have you ever wondered what it means to have a…