Building a Natural Language Processing Pipeline with Python: A Hands-On Workshop for Named Entity Recognition

If you’ve ever needed to label your data for downstream AI, you know how tedious and error-prone manual tagging can be. Fortunately, with the help of Python and NLP tools like spaCy, you can automate this process and even train a model that understands your domain.

In this workshop, you’ll build a system that classifies Spanish customer names as either personal names or company names using Named Entity Recognition (NER). We’ll begin by testing spaCy’s pre-trained Spanish model, which works decently in many cases — but you’ll soon see that the out-of-the-box accuracy may not be good enough for your needs.

That’s why the second half of this workshop shows you how to train your own custom NER model that performs better on your specific data.

📍 Outline

  1. Generate Sample Data (via ChatGPT)
  2. Set Up Your Environment
  3. Install Required Libraries
  4. Create and Run Python Code
  5. Test Built-In NER with spaCy (with sample output)
  6. Train a Custom NER Model
  7. View Improved Output
  8. Expand Your Dataset and Retrain
  9. Conclusion

1. Generate Sample Data

No need to download anything from outside websites. Just ask ChatGPT to generate your data files.

✨ To generate a sample Excel file:

Prompt to ChatGPT:
“Create an Excel file with one column titled ‘Customer Name’. Fill it with 50 rows — some with Spanish personal names (like ‘Juan Pérez’) and others with company names (like ‘ACME S.A.’). Send me the file for download.”

Save it as customer_names.xlsx into your project folder.

✨ To generate labeled training data for NER:

Prompt to ChatGPT:
“Generate a JSON list of short Spanish texts. Annotate each with start and end character indexes and label them as either ‘COMPANY’ or ‘PERSON’. Format it for spaCy NER training and send the file.”

Save the file as training_data.json.

2. Set Up Your Environment

Step 1: Install Python

Go to python.org and install Python 3. Make sure to check “Add Python to PATH” during setup.

Step 2: Install VS Code

Download and install Visual Studio Code: code.visualstudio.com

Step 3: Create a Project Folder

  1. Open VS Code
  2. Create a folder like ner_project
  3. Inside it, create a file called main.py

Step 4: How to Run Python Code in VS Code

  1. Open main.py in the editor.
  2. Use the Terminal → New Terminal option in VS Code to open a command line window.
  3. In the terminal, type:python main.py
  4. You’ll see your output directly in the terminal.

3. Install Required Libraries

In the VS Code terminal, run:

pip install pandas openpyxl spacy
python -m spacy download es_core_news_sm

4. Create and Run Python Code

Paste the following into main.py to load and preview the names:

import pandas as pd

# Load the Excel file
df = pd.read_excel("customer_names.xlsx")

# Select column B (index 1)
names = df.iloc[:, 1]

# Preview names
print("Sample names:")
print(names.head())

▶️ Run It:

python main.py

5. Test Built-In NER with spaCy (Baseline)

Now let’s run spaCy’s built-in Spanish NER model and inspect how it classifies each name:

import spacy

nlp = spacy.load("es_core_news_sm")

# Test first 10 names
print("\n--- Pre-trained spaCy Model Output ---")
for name in names.head(10):
    doc = nlp(name)
    if doc.ents:
        for ent in doc.ents:
            print(f"Name: {name} | Entity: {ent.text} | Label: {ent.label_}")
    else:
        print(f"Name: {name} | Entity: None")

📜 Sample Output (Before Training)

--- Pre-trained spaCy Model Output ---
Name: MARÍA EUGENIA LARA | Entity: MARÍA EUGENIA | Label: PER
Name: TRANSPORTES MARTÍNEZ S.A. | Entity: None
Name: JUAN CARLOS SALGADO | Entity: JUAN CARLOS | Label: PER
Name: INVERSIONES RÍO SUR LTDA. | Entity: RÍO | Label: LOC
Name: ANA LUCÍA GONZÁLEZ | Entity: ANA LUCÍA | Label: PER
Name: SERVICIOS ELECTROMECÁNICOS DEL NORTE | Entity: None

🔎 Analysis

As you can see, the pre-trained model can identify many PERSON names but struggles with COMPANY names. This is expected because general-purpose models often don’t understand industry-specific or regional company naming conventions.

That’s why we need to train our own NER model.

6. Train a Custom NER Model

Make sure you’ve saved training_data.json. It should look like:

[
  {
    "text": "MARÍA EUGENIA LARA",
    "spans": [{"start": 0, "end": 19, "label": "PERSON"}]
  },
  {
    "text": "SERVICIOS ELECTROMECÁNICOS DEL NORTE",
    "spans": [{"start": 0, "end": 39, "label": "COMPANY"}]
  }
]

Add the Following to main.py (or a separate script):

import json
from spacy.training.example import Example

# Load training data
with open("training_data.json", "r", encoding="utf-8") as f:
    training_data = json.load(f)

# Load model
nlp = spacy.load("es_core_news_sm")
ner = nlp.get_pipe("ner")

# Add labels
for entry in training_data:
    for span in entry["spans"]:
        ner.add_label(span["label"])

# Format examples
examples = []
for entry in training_data:
    doc = nlp.make_doc(entry["text"])
    entities = [(span["start"], span["end"], span["label"]) for span in entry["spans"]]
    examples.append(Example.from_dict(doc, {"entities": entities}))

# Train the model
optimizer = nlp.resume_training()
for i in range(10):  # 10 epochs
    losses = {}
    nlp.update(examples, drop=0.2, losses=losses)
    print(f"Epoch {i+1} Losses: {losses}")

# Save model
nlp.to_disk("company_person_ner")

7. View Improved Output (After Custom NER)

Once training is complete, test the new model:

# Load custom model
nlp_custom = spacy.load("company_person_ner")

print("\n--- Custom NER Model Output ---")
for name in names.head(10):
    doc = nlp_custom(name)
    if doc.ents:
        for ent in doc.ents:
            print(f"Name: {name} | Entity: {ent.text} | Label: {ent.label_}")
    else:
        print(f"Name: {name} | Entity: None")

📜 Sample Output (After Training)

--- Custom NER Model Output ---
Name: MARÍA EUGENIA LARA | Entity: MARÍA EUGENIA LARA | Label: PERSON
Name: TRANSPORTES MARTÍNEZ S.A. | Entity: TRANSPORTES MARTÍNEZ S.A. | Label: COMPANY
Name: JUAN CARLOS SALGADO | Entity: JUAN CARLOS SALGADO | Label: PERSON
Name: INVERSIONES RÍO SUR LTDA. | Entity: INVERSIONES RÍO SUR LTDA. | Label: COMPANY
Name: ANA LUCÍA GONZÁLEZ | Entity: ANA LUCÍA GONZÁLEZ | Label: PERSON
Name: SERVICIOS ELECTROMECÁNICOS DEL NORTE | Entity: SERVICIOS ELECTROMECÁNICOS DEL NORTE | Label: COMPANY

🔎 Analysis

Now, the model correctly identifies full names and even complex company names that were previously missed. Your model is now domain-aware!

8. Expand Your Dataset and Retrain

Now that you’ve seen how the model improves with a small labeled dataset, here’s your next move:

🛠️ Edit Your Excel and JSON Files

  1. Add More Names to your customer_names.xlsx file
    • Include edge cases like:
      • Abbreviated names (e.g., “M. González”)
      • Foreign names
      • Companies with common words (e.g., “Grupo Médico”)
  2. Expand Your Training Data in training_data.json
    • Make sure each new entry includes:{ "text": "NEW NAME OR COMPANY", "spans": [{"start": 0, "end": LENGTH, "label": "PERSON" or "COMPANY"}] }

💡 Tip: You can use ChatGPT again to generate a batch of 20–30 labeled entries formatted for spaCy training.

↻ Retrain Your Model

Once you’ve updated the JSON file:

  • Rerun the training section of your Python script
  • Then test the model again on your updated Excel file

Each time you do this, your model becomes more robust and better suited to your specific domain.

9. Conclusion

In this tutorial, you:

  • Created realistic data using ChatGPT
  • Set up your Python development environment
  • Explored built-in NLP tools using spaCy
  • Found their limitations on domain-specific data
  • Trained a custom NER model to fix the problem
  • Observed the improvement with real output
  • Expanded your dataset and retrained your model for even better results

You’ve now experienced what real-world NLP projects are like: part data prep, part testing, part training, and all within your control.

You can absolutely do this.