Skip to main content

Understanding Pandas DataFrames - Basics

Calculating read time…

Understanding Pandas DataFrames

Think of a Pandas DataFrame like a smart table (like Excel, but in Python). It has rows and columns where you can store numbers, names, dates — anything! The best part is You can do powerful operations with just simple code. Let's learn step by step! 🐼

What is a DataFrame?

A DataFrame is a 2-dimensional table

  • Rows = Each record (like each student)
  • Columns = Different features (like Name, Age, Marks)
  • Each column can store different types of data (numbers, text, etc.)

💡 Think of it like: A spreadsheet where you can do math, sorting, filtering — all with code!

Creating DataFrames - Two Easy Ways

Way 1: From a Dictionary (Most Common)

Keys become column names, values become the data:

import pandas as pd

data = {
    'Name': ['Alice', 'Bob', 'Charlie', 'David'],
    'Age': [25, 30, 35, 28],
    'City': ['NYC', 'LA', 'Chicago', 'Miami'],
    'Salary': [50000, 60000, 70000, 55000]
}

df = pd.DataFrame(data)
print(df)

Output:

      Name  Age     City  Salary
0    Alice   25      NYC   50000
1      Bob   30       LA   60000
2  Charlie   35  Chicago   70000
3    David   28    Miami   55000

Clean table with column names on top! 🎯

Way 2: From List of Lists

When your data is in lists, just specify column names separately:

data = [
    ['Alice', 25, 'NYC', 50000],
    ['Bob', 30, 'LA', 60000],
    ['Charlie', 35, 'Chicago', 70000]
]

columns = ['Name', 'Age', 'City', 'Salary']
df = pd.DataFrame(data, columns=columns)

Same result! Use whichever fits your data.

Real Example: Student Grades

Let's create a student grade sheet and learn operations one by one!

Step 1: Create the Student Data

import pandas as pd
import numpy as np

data = {
    'Student': ['Alice', 'Bob', 'Charlie', 'David', 'Eva'],
    'Math': [85, 90, 78, 92, 88],
    'Science': [88, 76, 92, 85, 90],
    'English': [92, 85, 88, 79, 94],
    'Attendance': [95, 87, 92, 88, 96]
}

df = pd.DataFrame(data)
print(df)

Output:

   Student  Math  Science  English  Attendance
0    Alice    85       88       92          95
1      Bob    90       76       85          87
2  Charlie    78       92       88          92
3    David    92       85       79          88
4      Eva    88       90       94          96

Perfect! Now we have 5 students with their marks. Let's analyze this data! 📊

Step 2: Add Total Marks Column

Want to calculate total? Just add columns together:

df['Total Marks'] = df['Math'] + df['Science'] + df['English']
print(df)

Output:

   Student  Math  Science  English  Attendance  Total Marks
0    Alice    85       88       92          95          265
1      Bob    90       76       85          87          251
2  Charlie    78       92       88          92          258
3    David    92       85       79          88          256
4      Eva    88       90       94          96          272

Magic! Pandas added all three subjects for each student automatically. 🎉

Step 3: Calculate Average Marks

Divide total by 3 to get average:

df['Average'] = df['Total Marks'] / 3
print(df)

Output:

   Student  Math  Science  English  Attendance  Total Marks    Average
0    Alice    85       88       92          95          265  88.333333
1      Bob    90       76       85          87          251  83.666667
2  Charlie    78       92       88          92          258  86.000000
3    David    92       85       79          88          256  85.333333
4      Eva    88       90       94          96          272  90.666667

Now we can see each student's average! But let's assign grades to make it clearer.

Assigning Grades - Two Methods

Method 1: Using NumPy (Fast Way)

When you have simple if-else conditions, use NumPy's select():

conditions = [
    df['Average'] >= 90,   # If average >= 90, give 'A'
    df['Average'] >= 80,   # If average >= 80, give 'B'
    df['Average'] >= 70    # If average >= 70, give 'C'
]

choices = ['A', 'B', 'C']

df['Grade'] = np.select(conditions, choices, default='D')
print(df)

Output:

   Student  Math  Science  English  Attendance  Total Marks    Average Grade
0    Alice    85       88       92          95          265  88.333333     B
1      Bob    90       76       85          87          251  83.666667     B
2  Charlie    78       92       88          92          258  86.000000     B
3    David    92       85       79          88          256  85.333333     B
4      Eva    88       90       94          96          272  90.666667     A

How it works:

  • Checks each condition from top to bottom
  • When a condition is True, picks the matching grade
  • If nothing matches, uses default ('D')

Only Eva scored 'A'!

Method 2: Using Pandas apply() (Flexible Way)

When you need more control or complex logic, use apply() with a custom function:

def assign_grade(avg):
    if avg >= 90:
        return 'A'
    elif avg >= 80:
        return 'B'
    elif avg >= 70:
        return 'C'
    else:
        return 'D'

df['Grade_Apply'] = df['Average'].apply(assign_grade)
print(df)

How it works:

  • Takes each value from 'Average' column one by one
  • Passes it to your function
  • Returns the grade and creates a new column

When to use which?

  • Use np.select() → Fast, simple conditions, great for big data
  • Use apply() → Need complex logic, multiple inputs, custom calculations

Filtering Students

Want to find only top students (Grade A)? Filter the data:

# Create a True/False column
df['Is_Top_Student'] = df['Grade'] == 'A'
print(df)

Or get only Grade A students directly:

top_students = df[df['Grade'] == 'A']
print(top_students)

This creates a new table with only 'A' grade students! Perfect for finding star performers.

Sorting Students by Marks

Let's arrange students from lowest to highest marks:

df = df.sort_values('Total Marks', ascending=True)
print(df)

Output:

   Student  Math  Science  English  Attendance  Total Marks    Average Grade
1      Bob    90       76       85          87          251  83.666667     B
3    David    92       85       79          88          256  85.333333     B
2  Charlie    78       92       88          92          258  86.000000     B
0    Alice    85       88       92          95          265  88.333333     B
4      Eva    88       90       94          96          272  90.666667     A

Now Bob is at the bottom (lowest marks) and Eva at the top (highest marks)!

Want highest first? Just change ascending=False:

df = df.sort_values('Total Marks', ascending=False)

Adding Rank

Let's give each student a rank based on their total marks:

df['Rank'] = df['Total Marks'].rank(ascending=False)
print(df)

Output:

   Student  Math  Science  English  Attendance  Total Marks    Average Grade  Rank
0    Alice    85       88       92          95          265  88.333333     B   2.0
1      Bob    90       76       85          87          251  83.666667     B   5.0
2  Charlie    78       92       88          92          258  86.000000     B   3.0
3    David    92       85       79          88          256  85.333333     B   4.0
4      Eva    88       90       94          96          272  90.666667     A   1.0

How ranking works:

  • ascending=False means highest marks get Rank 1
  • ascending=True would give lowest marks Rank 1
  • Eva got Rank 1 (highest marks), Bob got Rank 5 (lowest marks)

Perfect! Now you can see who's performing best at a glance.

Quick Summary 📝

What we learned today:

  • Creating DataFrames → From dictionary or list of lists
  • Adding columns → Simple math operations like df['New'] = df['A'] + df['B']
  • Conditional logic → Using np.select() (fast) or apply() (flexible)
  • Filtering → Get specific rows with df[df['Column'] == value]
  • Sorting → Arrange data with sort_values()
  • Ranking → Assign ranks with rank()

Keep practicing with your own data!. Happy learning! 🐼

Comments