Tutorial
Entity extraction using Watson NLP
Step through the processing of using Watson NLP to extract entities, keywords, and phrasesArchive date: 2025-12-19
This content is no longer being updated or maintained. The content is provided “as is.” Given the rapid evolution of technology, some content, steps, or illustrations may have changed.Entity, keyword, and phrase extraction plays a major role in understanding unstructured data. These entities can include names of people, organization names, dates, prices, and facilities. IBM Watson NLP now provides the ability to automatically extract entities, keywords, and phrases using pretrained models.
IBM Watson NLP is a standard embeddable AI library that is designed to tie together the pieces of IBM Natural Language Processing. It provides a standard base natural language processing layer along with a single integrated roadmap, a common architecture, and a common code stack designed for widespread adoption across IBM products.
The watson_nlp library is available on IBM Watson Studio as a runtime library so that you can directly use it for model training, evaluation, and prediction. The following figure shows the Watson NLP architecture.

This tutorial explains the fundamentals of IBM Watson NLP and walks you through the process of using it to extract entities, keywords, and phrases.
Prerequisites
To follow the steps in this tutorial, you need:
- An IBMid
- A Watson Studio project
- A Python notebook
- Your environment set up
Estimated time
It should take you approximately 1 hour to complete this tutorial.
Steps
The steps in this tutorial use examples of hotel reviews from Kaggle 515 K Hotel Reviews Data in Europe and OpinRank Review Data set to walk you through the process.
Step 1. Collecting the data set
Note: If you are reserving the environment through the IBM Tech Zone, you don't need to collect the data manually. The environment comes with the Watson Studio project pre-created for you. You can skip the rest of the steps here and follow the instructions in the notebook to complete the Text Classification tutorial.
If you are not reserving the environment through the Tech Zone and you have a Watson Studio instance, then you should use the following steps.
Download the data set from three hotels for analysis and comparison.
Upload the data set to your Watson Studio project by going to the Assets tab, and then dropping the data files as shown in the following figure.

After you have added the data set to the project, you might have to reload the Notebook. You have two options of accessing the data set from the Jupyter Notebook, depending on the level of access you have.
A. If you are a project administrator, then:
i) You can just insert the project token as shown in the following image.

ii) After inserting the project token, you can continue running all of the cells in the notebook. This cell in particular loads your data set in the notebook.

B. If you are not a Watson Studio project administrator, then you cannot create a project token.
i) Create a new cell under Step 2 - Data Loading by clicking the Insert menu, and then selecting Insert Cell Below or the Esc+B keyboard shortcut. Highlight the code cell that is shown in the following image by clicking it.

ii) Ensure that you place the cursor below the commented line. Click the Find and add data icon (01/00) in the upper right.
iii) Choose the Files tab, and pick the
uk_england_london_belgrave_hotel.csvfile. Click Insert to code, and choose pandas DataFrame. Rename the DataFrame fromdf_data_1tobelgrave_df.iv) Choose the Files tab, and pick the
uk_england_london_euston_square_hotel.csvfile. Click Insert to code, and choose pandas DataFrame. Rename the DataFrame fromdf_data_2toeuston_df.v) Choose the Files tab, and pick the
uk_england_london_dorset_square.csvfile. Click Insert to code, and choose pandas DataFrame. Rename the DataFrame fromdf_data_3todorset_df.vi) Choose the Files tab, and pick the
london_hotel_reviews.csvfile. Click Insert to code, and choose pandas DataFrame. Rename the DataFrame fromdf_data_4tohotels_df.
After you've added the data set to the project, you can access it from the Jupyter Notebook, and read the .csv file into a pandas DataFrame.

Step 3. Entity extraction
Entity extraction uses the entity-mentions block to encapsulate algorithms for the task of extracting mentions of entities (persons, organizations, dates, and locations) from the input text. The block offers implementations of strong entity extraction algorithms from each of the four families: rule-based, classic machine learning, deep-learning, and transformers.
There are two types of models:
- A rule-based model (the rbr model), which handles syntactically regular entity types such as number, email, and phone.
- A model that is trained on labeled data for the more complex entity types such as persons, organizations, and locations.
Step 3.1. Entity extraction function
Rule-based models (rbr) do not depend on any blocks, so you can directly run them on input text.
Models that are trained from labeled data, such as BilSTM, BERT, and transformer depend on the syntax block. As such, the syntax block must be run first to generate the input expected by the entity-mention block.
Load the syntax model and three entity extraction models.
# Load a syntax model to split the text into sentences and tokens syntax_model = watson_nlp.load(watson_nlp.download('syntax_izumo_en_stock')) # Load bilstm model in WatsonNLP bilstm_model = watson_nlp.load(watson_nlp.download('entity-mentions_bilstm_en_stock')) # Load rbr model in WatsonNLP rbr_model = watson_nlp.load(watson_nlp.download('entity-mentions_rbr_en_stock')) # Load bert model in WatsonNLP bert_model = watson_nlp.load(watson_nlp.download('entity-mentions_bert_multi_stock'))Build a custom function to run a specified entity extraction model and parse its results. The returned output is a dictionary of the review text, hotel name, website, and entity mentions.
def extract_entities(data, model, hotel_name=None, website=None): import html input_text = str(data) text = html.unescape(input_text) if model == 'rbr': # Run rbr model on text mentions = rbr_model.run(text) else: # Run syntax model on text syntax_result = syntax_model.run(text) if model == 'bilstm': # Run bilstm model on syntax result mentions = bilstm_model.run(syntax_result) elif model == 'bert': # Run bert model on syntax result mentions = bert_model.run(syntax_result) elif model == 'transformer': # Run transformer model on syntax result mentions = transformer_model.run(syntax_result) entities_list = mentions.to_dict()['mentions'] print(entities_list) ent_list=[] for i in range(len(entities_list)): ent_type = entities_list[i]['type'] ent_text = entities_list[i]['span']['text'] ent_list.append({'ent_type':ent_type,'ent_text':ent_text}) if len(ent_list) > 0: return {'Document':input_text,'Hotel Name':hotel_name,'Website':website,'Entities':ent_list} else: return {}Stop-words are common words that are not meaningful for separating the data. Such common words are assumed to be "noise" because their high frequency might hide the words carrying more informative signals. Here, these are filtered based on a predefined list that is used in Watson NLP and based on the part-of-speech.
Note: The stop-words list can be customized for the target data set. This is demonstrated in the following code example. When the documents are vectorized in the following code, a filter is applied that ignores terms that appear in 50% or more of the documents. This filter can also be counted as part of stop-words filtering.
wnlp_stop_words = watson_nlp.download_and_load('text_stopwords_classification_ensemble_en_stock').stopwords stop_words = list(wnlp_stop_words) stop_words.remove('keep') stop_words.extend(["gimme", "lemme", "cause", "'cuz", "imma", "gonna", "wanna", "gotta", "hafta", "woulda", "coulda", "shoulda", "howdy","day", "first", "second", "third", "fourth", "fifth", "London", "london", "1st", "2nd", "3rd", "4th", "5th", "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday", "weekend", "week", "evening", "morning"])
Step 3.2. Run entity extraction
Apply data preprocessing on the input text or documents, and then run the cleaned text through the model. The model used can be specified through the model parameter in
extract_entities().def run_extraction(df_list, text_col): extract_list = [] for df in df_list: all_text = dict(zip(df[text_col], zip(df['hotel'], df['website']))) all_text_clean = {clean(doc[0]): doc[1] for doc in all_text.items()} for text in all_text_clean.items(): # change the second parameter to 'rbr', 'bilstm', or 'bert' to try other models extract_value = extract_entities(text[0], 'bilstm', text[1][0], text[1][1]) if len(extract_value) > 0: extract_list.append(extract_value) return extract_listThe model outputs a text's entity mention as well as its category of entity. For example, a "London" mention is a Location type, and "good soundproof rooms" is a Facility type.

The Entities extraction model output can be visualized to analyze and understand the entities in the hotel reviews clearly.

Step 3.3. Comparing top 20 entities for each hotel
You can examine the results of the entity extraction by plotting the top frequently mentioned entities for each hotel. These mentions can be used to generate tags for a hotel to create relevancy and familiarity for search engine results. The following figure shows the display frequency with horizontal bar charts.

Step 3.4. Comparison between Booking.com and TripAdvisor for one hotel
Another approach to analyzing hotel customer reviews is by comparing the entity mentions found on the two websites where the reviews are published. This example looks at Booking.com versus TripAdvisor to try to gain insight on the tendencies of reviewers who use one platform compared to the other platform.
Plot side-by-side word clouds for each of the hotels

You see that the customers on Booking.com care more about the convenience to Tube stations whereas TripAdvisor customers care more about the reception and atmosphere of the hotel itself.
Look at the Euston hotel word clouds.

You can use this collective information to give priority to the website with reviews that better align with your preferences about choosing a hotel. Do you care more about the convenience of the location of a hotel or do you care about the hotel's ambience, reception, or perks?
Step 4. Keyword phrase extraction
Another Watson NLP capability that you can use to analyze the hotel reviews is keyword phrase extraction. The keywords block ranks noun phrases that are extracted from an input document based on how relevant they are within the document.
This tutorial uses the text-rank model. The text-rank model takes the output of noun phrase models and assigns a relevance score for each extracted noun phrase. The relevance score calculation is inspired by the page rank algorithm. In the context of the input document, extracted noun phrases that appear in “more connected” contexts receive a higher rank. Additionally, the relevance score of extracted noun phrases is upgraded when the noun phrase appears more frequently in Wikipedia.
Load noun phrases, embedding, and keywords models for English.
syntax_model = watson_nlp.load(watson_nlp.download('syntax_izumo_en_stock')) noun_phrases_model = watson_nlp.load(watson_nlp.download('noun-phrases_rbr_en_stock')) keywords_model = watson_nlp.load(watson_nlp.download('keywords_text-rank_en_stock'))The following function captures the flow for running a keyword extraction model. First, the input document/text is run through the syntax model and the noun phrases model to extract noun texts. Then, both the syntax output and the noun output are used as inputs for the keyword model where the output includes the text phrase and relevance score. You can use this score to rank the phrases.
def extract_keywords(text): # Run the Syntax and Noun Phrases models syntax_prediction = syntax_model.run(text, parsers=('token', 'lemma', 'part_of_speech')) noun_phrases = noun_phrases_model.run(text) # Run the keywords model keywords = keywords_model.run(syntax_prediction, noun_phrases, limit=5) keywords_list = keywords.to_dict()['keywords'] key_list = [] for i in range(len(keywords_list)): dict_list = {} key = custom_tokenizer(keywords_list[i]['text']) dict_list['phrase'] = key dict_list['relevance'] = keywords_list[i]['relevance'] key_list.append(dict_list) return {'Complaint data':text,'Phrases':key_list}Plot the result of the keyword phrase extraction for the three hotels.

You can use these top-ranked phrases to generate a brief description of the most notable attributes of each hotel. Customers can find interest in a particular hotel with just one look at the list of phrases.
Step 5. Entity and phrase search
One of the applications of entity detection is in searching. You can search for entity mentions and phrases by using them as tags of reviews in the large London hotel reviews data set.
Sample 5% of the London hotels reviews for the sake of runtime and capacity.
hotels_df_sample = hotels_df.sample(frac = 0.05, random_state = 1)Run entity extraction on the sampled DataFrame.
hotels_extract_list = run_extraction([hotels_df_sample], 'text')Build a function to run and parse phrases on the sampled DataFrame. This function uses the previously built
extract_keywords()function.def explode_phrases2(hotels_df): keywords = [] for index, row in hotels_df.iterrows(): keywords.append(extract_keywords(row['Document'], row['Hotel Name'])) phrases_df = pd.DataFrame(keywords) exp_phrases = phrases_df.explode('Phrases') exp_phrases = exp_phrases.dropna(subset=['Phrases']) exp_phrases = pd.concat([exp_phrases.drop(['Phrases'], axis=1), exp_phrases['Phrases'].apply(pd.Series)], axis=1) exp_phrases['phrase_length'] = exp_phrases['phrase'].apply(lambda x: len(x.split(' '))) # Removing uni-gram and bi-grams exp_phrases = exp_phrases[exp_phrases.phrase_length > 2] return exp_phrasesCombine the results of the entity extraction and phrase extraction functions into one DataFrame.
Input a list of entities and a list of phrases to the search function and run it on the DataFrame.
def search_entity(hotels_df, entities_list, phrases_list): search_df = hotels_df[(hotels_df['ent_text'].str.lower().str.contains('|'.join(entities_list).lower())) & (hotels_df['phrase'].str.lower().str.contains('|'.join(phrases_list).lower()))] hotel_count = search_df['Hotel Name'].value_counts().to_dict() return search_df, hotel_countFor example, if you care about the food, cleanliness, and convenience of a hotel, you can use the words
good breakfast,room service,clean, andshoppingto look for hotel reviews that mention these entities and phrases. You hope to find a short-end list of hotels that fulfills your criteria.
Based on the results of the entity and phrase extraction search, you can determine that these five hotels best offer the features that you are looking for. To get a better idea of how the hotels were received, you can read all of the reviews from each of the hotels that matched.
Step 6. Actionable insights using entities and keyword phrase extraction combined with sentiment analysis
For the five hotels, you can determine the targeted sentiment for the phrases that are found in their reviews. You want to ensure that the entities and phrases you've detected are positive so that you can make an accurate decision on which hotel to pick.
The Targets Sentiment Extraction block contains algorithms for the task of extracting the sentiments that are expressed in text and identifying the targets of those sentiments. The block automatically outputs both the target terms and the sentiment that is expressed toward each target term when given an input text.
For example, given the input:
“The served food was delicious, yet the service was slow.”
The block identifies that there is a positive sentiment expressed toward the target “food”, and a negative sentiment expressed toward “service”. A significant advantage of this block is that it can handle multiple targets with different sentiments in one sentence.
Load the target sentiment extraction model.
sentiment_extraction_model = watson_nlp.load(watson_nlp.download('targets-sentiment_sequence-bert_multi_stock'))Build a function to run the sentiment extraction model on the entity and phrase results from the previous step.
def run_sentiment(df, text_col, ent_col): pos_targets =[] neg_targets =[] targets = [] entities = dict(df[ent_col]) for text, hotel in zip(df[text_col], df['Hotel Name']): syntax_analysis_en = syntax_model.run(text, parsers=('token',)) extracted_sentiments = sentiment_extraction_model.run(syntax_analysis_en) for key , score in extracted_sentiments.to_dict()['targeted_sentiments'].items(): label = score['label'] targets.append({'Hotel Name' : hotel, 'phrase' : key, 'sentiment' : score['label']}) if label=='SENT_POSITIVE': # and key not in pos_targets: pos_targets.append(key) elif label=='SENT_NEGATIVE': # and key not in neg_targets: neg_targets.append(key) return pos_targets, neg_targets, targetsThe function captures the flow for running a target sentiment extraction model. First, you run the input document/text through the syntax model. Then, you use the syntax output as input for the target sentiment extraction model where the output includes the target text and their sentiment. You return the outputs into two separate lists of positive and negative sentiment text as well as one list of labels for every phrase.
You can display a word cloud for the positive sentiment phrases of every hotel in the search result to get a better view at the well-rated features of each hotel.

Conclusion
This tutorial showed how to use the Watson NLP library and how easily you can run various entity, phrase, and target sentiment extraction models on input text. This notebook also demonstrated one possible application of Watson NLP.
Work through the notebook to try out this feature.