{"id":111620,"date":"2019-04-10T15:48:00","date_gmt":"2019-04-10T20:48:00","guid":{"rendered":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/decoded\/\/\/overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach\/"},"modified":"2024-04-14T04:10:40","modified_gmt":"2024-04-14T09:10:40","slug":"overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach","status":"publish","type":"decoded","link":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/decoded\/2019\/04\/10\/overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach\/","title":{"rendered":"Overcoming the limitations of topic models with a semi-supervised approach"},"content":{"rendered":"\n<figure class=\"wp-block-image size-640-wide\"><a rel=\"attachment wp-att-125986\" href=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/decoded\/\/\/overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach\/04-10-2019_feature-png\/\"><img data-dominant-color=\"efefef\" data-has-transparency=\"false\" style=\"--dominant-color: #efefef;\" loading=\"lazy\" decoding=\"async\"  srcset=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?resize=480,270 480w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?resize=700,394 700w\" sizes=\"(max-width: 480px) 480px, (max-width: 782px) 782px, 640px\" height=\"360\" width=\"640\" src=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=640\" alt=\"\" class=\"wp-image-125986 not-transparent\" \/><\/a><figcaption class=\"wp-element-caption\"><em>(Related posts:&nbsp;<\/em><a href=\"https:\/\/medium.com\/pew-research-center-decoded\/an-intro-to-topic-models-for-text-analysis-de5aa3e72bdb\"><em>An intro to topic models for text analysis<\/em><\/a><em>,&nbsp;<\/em><a href=\"https:\/\/medium.com\/pew-research-center-decoded\/making-sense-of-topic-models-953a5e42854e\"><em>Making sense of topic models<\/em><\/a>,&nbsp;<a href=\"https:\/\/medium.com\/pew-research-center-decoded\/interpreting-and-validating-topic-models-ff8f67e07a32\"><em>Interpreting and validating topic models<\/em><\/a><em>,&nbsp;<\/em><a href=\"https:\/\/medium.com\/pew-research-center-decoded\/how-keyword-oversampling-can-help-with-text-analysis-c15c9c410c0c\"><em>How keyword oversampling can help with text analysis<\/em><\/a><em>&nbsp;and&nbsp;<\/em><a href=\"https:\/\/medium.com\/pew-research-center-decoded\/are-topic-models-reliable-or-useful-c960f945c9cb\"><em>Are topic models reliable or useful?<\/em><\/a><em>)<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"df66\">In two earlier posts on this blog, I introduced&nbsp;<a href=\"https:\/\/medium.com\/pew-research-center-decoded\/an-intro-to-topic-models-for-text-analysis-de5aa3e72bdb\">topic models<\/a>&nbsp;and explored&nbsp;<a href=\"https:\/\/medium.com\/pew-research-center-decoded\/making-sense-of-topic-models-953a5e42854e\">some of the difficulties<\/a>&nbsp;that can arise when researchers attempt to use them to measure content. In this post, I\u2019ll show how to overcome some of these challenges with what\u2019s known as a \u201csemi-supervised\u201d approach.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"f138\">To illustrate how this approach works, I\u2019ll use a dataset of open-ended survey responses to the following question: \u201cWhat makes your life feel meaningful, satisfying or fulfilling?\u201d The dataset comes from a recent Pew Research Center report about&nbsp;<a href=\"http:\/\/alpha.pewresearch.org\/pewresearch-org\/religion\/2018\/11\/20\/where-americans-find-meaning-in-life\/\" target=\"_blank\" rel=\"noreferrer noopener\">where Americans find meaning in their lives<\/a>. (You can read some of the actual responses&nbsp;<a href=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/religion\/interactives\/what-keeps-us-going\/\" target=\"_blank\" rel=\"noreferrer noopener\">here<\/a>.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As my colleagues and I worked on this project over the past year, we tested a number of computational methods \u2014 including several topic modeling algorithms \u2014 to measure the themes that emerged from these survey responses. The image below shows a selection of four topics drawn from models trained on our dataset (the topics are shown in the rows). To arrive at these topics, we used three different topic modeling algorithms, two of which are shown in the columns below \u2014 Latent Dirichlet Allocation (LDA) and Non-Negative Matrix Factorization (NMF). (More on the third algorithm later.)<\/p>\n\n\n\n<figure class=\"wp-block-image size-640-wide\"><a rel=\"attachment wp-att-125988\" href=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/decoded\/\/\/overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach\/image-5-png-2\/\"><img data-dominant-color=\"f2ebea\" data-has-transparency=\"false\" style=\"--dominant-color: #f2ebea;\" loading=\"lazy\" decoding=\"async\"  srcset=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-5.png?resize=480,275 480w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-5.png?resize=782,448 782w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-5.png?resize=960,550 960w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-5.png?resize=1200,687 1200w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-5.png?resize=1400,802 1400w\" sizes=\"(max-width: 480px) 480px, (max-width: 782px) 782px, 640px\" height=\"367\" width=\"640\" src=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-5.png?w=640\" alt=\"\" class=\"wp-image-125988 not-transparent\" \/><\/a><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"8585\">In each of these models, we can clearly see topics related to four distinct themes: health, finances, career and religion (specifically Christianity). But it\u2019s also clear that many of these topics suffer from a few of the problems that I discussed in my&nbsp;<a href=\"https:\/\/medium.com\/pew-research-center-decoded\/making-sense-of-topic-models-953a5e42854e\">last post<\/a>: Some topics appear to be \u201covercooked\u201d or \u201cundercooked,\u201d and some contain what might be called \u201cconceptually spurious\u201d words, highlighted above in red. These are words that might apply across multiple contexts and can cause problems under some circumstances.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"f769\">For example, if our survey respondents only ever mention the word \u201cinsurance\u201d when discussing health insurance, it might still fit with a topic intended to measure the concept of \u201chealth.\u201d On the other hand, including the word \u201ccare\u201d in such a topic might be particularly problematic if, in addition to talking about \u201chealth care,\u201d our respondents also frequently use the word to talk about other themes unrelated to health, like how much they care about their families. If that were the case, our model might substantially overestimate the prevalence of the true topic of \u201chealth\u201d in our dataset.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"dcc5\">For these reasons, our initial attempts at topic modeling Americans\u2019 sources of meaning in life raised some concerns \u2014 while we were confident that we could use this method to identify the main themes in our documents, it was unclear whether it could&nbsp;<em>reliably measure<\/em>&nbsp;those themes. While topic models can be executed quickly, they don\u2019t always make classification decisions as accurately as more complex supervised learning models would \u2014 and in some cases their results can be downright misleading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"3c84\">Fortunately, one recent innovation offered us a compromise between an unsupervised topic modeling approach and a more labor-intensive supervised classification modeling approach: semi-supervised topic models. While still largely automated, the new class of semi-supervised topic models allows researchers to draw on their domain knowledge to \u201cnudge\u201d the model in the right direction.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"df93\">Intrigued by this possibility, we tested a third algorithm called&nbsp;<a href=\"https:\/\/github.com\/gregversteeg\/corex_topic\">CorEx<\/a>&nbsp;(short for Correlation Explanation), a semi-supervised topic model that \u2014 unlike LDA and NMF \u2014 allowed us to provide the model with \u201canchor words\u201d that represented potential topics we thought the model should attempt to find.&nbsp;With CorEx, you can also tell the model how much weight it should give to these anchors. If you are less confident in your choices, the model may override your suggestions if they don\u2019t fit the data well enough. Armed with this new ability to guide a topic model with our domain expertise, we set out to assess whether we could make topic models more useful by manually correcting some of the problems we discovered in the above topics.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"ccc2\">In training our initial CorEx models, we embraced the reality that no perfect number of topics exists. No matter which parameters you set, a topic model of any size almost always returns at least a handful of jumbled or uninterpretable topics. We decided to train two different models \u2014 one with a large number of topics and one with fewer topics \u2014 and found that some of the topics that were \u201cundercooked\u201d in the smaller model were successfully split apart in the larger one.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"e017\">By the same token, \u201covercooked\u201d topics in the larger model were collapsed into more general and interpretable topics in the smaller one. We didn\u2019t use any anchor words in this first run because we wanted to use the models to identify the main themes in our data and avoid imposing our own expectations about the prominent topics in the responses.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"2be1\">We brought the anchors into play as we reviewed the topics produced by the two models. After getting a sense of the core themes the models were picking up, we selected words that seemed to correctly belong to each topic and \u201cconfirmed\u201d them by setting them as anchors. Then, in addition to drawing on our own domain knowledge to expand our lists of anchors with words we knew were related, we also read through a sample of our survey responses and added any relevant words that we noticed the models had missed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"1f85\">In cases where the models were incorrectly adding \u201cconceptually spurious\u201d words to certain topics, we encouraged the models to pull the undesirable words out of our good topics by adding them to a list of anchor words for a separate \u201cjunk topic\u201d that we had designated for this purpose. After re-running the models and repeating this process several times, our anchor lists had grown substantially, and our topics were looking much more interpretable and coherent:<\/p>\n\n\n\n<figure class=\"wp-block-image size-640-wide\"><a rel=\"attachment wp-att-125990\" href=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/decoded\/\/\/overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach\/image-6-png-2\/\"><img data-dominant-color=\"e9e9e9\" data-has-transparency=\"false\" style=\"--dominant-color: #e9e9e9;\" loading=\"lazy\" decoding=\"async\"  srcset=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-6.png?resize=480,263 480w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-6.png?resize=782,429 782w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-6.png?resize=960,527 960w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-6.png?resize=1200,658 1200w, https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-6.png?resize=1400,768 1400w\" sizes=\"(max-width: 480px) 480px, (max-width: 782px) 782px, 640px\" height=\"351\" width=\"640\" src=\"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/image-6.png?w=640\" alt=\"\" class=\"wp-image-125990 not-transparent\" \/><\/a><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"5bef\">Our application of semi-supervised topic modeling was a game-changer. By guiding the models in the right direction over multiple iterations and careful revisions of our anchor lists, we were able to successfully remove many of the conceptually spurious and multi-context words that appeared in the initial unanchored models. Our topics were looking much cleaner, and we were now hopeful that we could use this approach not only to identify the main topics in our data, but also refine them enough to interpret them, give them clear labels and use them as reliable measurement instruments. In my upcoming posts, I\u2019ll cover how we did just that.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" id=\"7296\">In the meantime, if you\u2019d like to try out semi-supervised topic modeling yourself, here is some&nbsp;<a href=\"https:\/\/gist.github.com\/patrickvankessel\/0d5bd690910edece831dbdf32fb2fb2d\" rel=\"noreferrer noopener\" target=\"_blank\">example code<\/a>&nbsp;to get you started.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In this post, I\u2019ll show how to overcome some of challenges that arise with topic modeling with what\u2019s known as a \u201csemi-supervised\u201d approach.<\/p>\n","protected":false},"author":655,"featured_media":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"sub_headline":"","sub_title":"","_prc_public_revisions":[],"_ppp_expiration_hours":0,"_ppp_enabled":false,"ai_generated_summary":"","relatedPosts":[],"datacite_doi":"","datacite_doi_citation":"","_prc_seo_qr_attachment_id":0,"spoken_article_player_enabled":true,"displayBylines":true,"footnotes":"","prc_watchers":[],"_prc_fork_parent":0,"_prc_fork_status":"","_prc_active_fork":0},"categories":[353],"bylines":[927],"collection":[],"_post_visibility":[],"decoded-category":[530],"formats":[],"_fund_pool":[],"languages":[],"regions-countries":[],"research-teams":[524],"workflow-status":[],"class_list":["post-111620","decoded","type-decoded","status-publish","hentry","category-data-science","bylines-patrick-van-kessel","decoded-category-data-science","research-teams-decoded"],"label":"Decoded","post_parent":0,"word_count":1131,"canonical_url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/decoded\/2019\/04\/10\/overcoming-the-limitations-of-topic-models-with-a-semi-supervised-approach\/","art_direction":{"A1":{"id":125986,"rawUrl":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png","url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=564&h=317&crop=1","width":564,"height":317,"caption":"","chartArt":false},"A2":{"id":125986,"rawUrl":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png","url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=268&h=151&crop=1","width":268,"height":151,"caption":"","chartArt":false},"A3":{"id":125986,"rawUrl":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png","url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=194&h=110&crop=1","width":194,"height":110,"caption":"","chartArt":false},"A4":{"id":125986,"rawUrl":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png","url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=268&h=151&crop=1","width":268,"height":151,"caption":"","chartArt":false},"XL":{"id":125986,"rawUrl":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png","url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=700&h=394&crop=1","width":700,"height":394,"caption":"","chartArt":false},"social":{"id":125986,"rawUrl":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png","url":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-content\/uploads\/sites\/20\/2022\/08\/04.10.2019_feature.png?w=700&h=394&crop=1","width":700,"height":394,"caption":"","chartArt":false}},"_embeds":[],"watchers":[],"table_of_contents":[],"datacite_doi":"","prc_seo_data":{"title":"Overcoming the limitations of topic models with a semi-supervised approach","description":"In this post, I\u2019ll show how to overcome some of challenges that arise with topic modeling with what\u2019s known as a \u201csemi-supervised\u201d approach.","og_title":"Overcoming the limitations of topic models with a semi-supervised approach","og_description":"In this post, I\u2019ll show how to overcome some of challenges that arise with topic modeling with what\u2019s known as a \u201csemi-supervised\u201d approach.","schema_type":"Article","noindex":false,"canonical_url":"","primary_terms":{"category":1},"custom_schema":[],"og_image":125986,"indexnow_submitted_at":null,"gsc_index_status":null},"prepublish_checks":{},"jetpack_sharing_enabled":true,"relatedPostsOrdered":[],"bylinesOrdered":[{"key":"_iaxqoo1c3","termId":927}],"acknowledgementsOrdered":[],"_links":{"self":[{"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/decoded\/111620","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/decoded"}],"about":[{"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/types\/decoded"}],"author":[{"embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/users\/655"}],"replies":[{"embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/comments?post=111620"}],"version-history":[{"count":2,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/decoded\/111620\/revisions"}],"predecessor-version":[{"id":138540,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/decoded\/111620\/revisions\/138540"}],"wp:attachment":[{"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/media?parent=111620"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/categories?post=111620"},{"taxonomy":"bylines","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/bylines?post=111620"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/collection?post=111620"},{"taxonomy":"_post_visibility","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/_post_visibility?post=111620"},{"taxonomy":"decoded-category","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/decoded-category?post=111620"},{"taxonomy":"formats","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/formats?post=111620"},{"taxonomy":"_fund_pool","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/_fund_pool?post=111620"},{"taxonomy":"languages","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/languages?post=111620"},{"taxonomy":"regions-countries","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/regions-countries?post=111620"},{"taxonomy":"research-teams","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/research-teams?post=111620"},{"taxonomy":"workflow-status","embeddable":true,"href":"https:\/\/alpha.pewresearch.org\/pewresearch-org\/wp-json\/wp\/v2\/workflow-status?post=111620"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}