Python 工具链
Python 机器学习工具链
Section titled “Python 机器学习工具链”Python 是 AI 开发的主流语言。机器学习的数据处理流程可以概括为:
flowchart LR A[原始数据] --> B[NumPy<br/>数值计算] B --> C[Pandas<br/>数据整理] C --> D[Matplotlib<br/>可视化] D --> E[建模]三个库各司其职:NumPy 负责底层数值运算,Pandas 负责表格化数据整理,Matplotlib 负责可视化。
NumPy:数值计算基石
Section titled “NumPy:数值计算基石”NumPy 的核心是 ndarray——多维数组。和 Python 列表的区别:
| 特性 | Python 列表 | NumPy 数组 |
|---|---|---|
| 元素类型 | 可以混合 | 必须统一 |
| 内存布局 | 分散存储 | 连续存储 |
| 运算速度 | 慢(循环) | 快(向量化) |
| 数学运算 | 需要手动写 | 内置支持 |
import numpy as np
# 创建数组的 6 种方式a = np.array([1, 2, 3, 4, 5]) # 从列表创建b = np.zeros((3, 4)) # 全零矩阵c = np.ones((2, 3)) # 全一矩阵d = np.eye(3) # 单位矩阵e = np.arange(0, 10, 2) # 等差数列 [0, 2, 4, 6, 8]f = np.linspace(0, 1, 5) # 等间距 5 个点g = np.random.randn(3, 3) # 标准正态分布随机矩阵
# 查看数组属性print(f"形状: {a.shape}, 维度: {a.ndim}, 元素类型: {a.dtype}")NumPy 的切片是视图而非拷贝——修改切片会影响原数组,这是性能优化的关键设计。
arr = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
print(arr[0, 0]) # 1 — 单个元素print(arr[0, :]) # [1, 2, 3] — 第一行print(arr[:, 1]) # [2, 5, 8] — 第二列print(arr[1:, :2]) # [[4,5],[7,8]] — 子矩阵(后两行,前两列)
# 布尔索引:筛选满足条件的元素mask = arr > 5print(arr[mask]) # [6, 7, 8, 9] — 所有大于 5 的元素广播是 NumPy 最强大的特性之一。当两个形状不同的数组运算时,NumPy 自动扩展较小的数组:
a = np.array([[1, 2, 3], [4, 5, 6]]) # 形状: (2, 3)b = np.array([10, 20, 30]) # 形状: (3,)
# b 自动"广播"为 (2, 3):# [[10, 20, 30],# [10, 20, 30]]print(a + b)# [[11, 22, 33],# [14, 25, 36]]广播规则:从最后一维开始对齐,维度为 1 或缺失的维度会被扩展。
x = np.array([1, 2, 3, 4, 5])
# 统计运算print(np.mean(x)) # 3.0 — 均值print(np.std(x)) # 1.41 — 标准差print(np.min(x)) # 1 — 最小值print(np.argmax(x)) # 4 — 最大值所在的索引
# 线性代数A = np.array([[1, 2], [3, 4]])B = np.array([[5, 6], [7, 8]])print(A @ B) # 矩阵乘法print(np.linalg.inv(A)) # 逆矩阵
# 变形print(x.reshape(5, 1)) # (5,) → (5, 1),增加一个维度print(x[:, np.newaxis]) # 同上,另一种写法Pandas:数据处理
Section titled “Pandas:数据处理”Pandas 的核心是 DataFrame——可以理解为”Python 中的 Excel 表格”。
flowchart TD A[CSV/Excel/数据库] --> B[DataFrame] B --> C[筛选 & 过滤] B --> D[分组 & 聚合] B --> E[清洗 & 转换] C --> F[分析结果] D --> F E --> F创建和查看数据
Section titled “创建和查看数据”import pandas as pd
df = pd.DataFrame({ "name": ["Alice", "Bob", "Charlie", "Diana"], "age": [25, 30, 35, 28], "salary": [50000, 60000, 80000, 55000], "dept": ["HR", "Eng", "Eng", "HR"],})
print(df.head(2)) # 查看前两行print(df.shape) # (4, 4) — 行数和列数print(df.dtypes) # 每列的数据类型print(df.describe()) # 数值列的统计摘要(均值、标准差、四分位数等)# 单条件:年龄大于 28print(df[df["age"] > 28])
# 多条件:年龄大于 28 且属于工程部print(df[(df["age"] > 28) & (df["dept"] == "Eng")])
# 按标签选择(loc)vs 按位置选择(iloc)print(df.loc[0:2, "name":"salary"]) # 按行标签和列标签print(df.iloc[0:2, 1:3]) # 按行位置和列位置# 按部门统计平均薪资print(df.groupby("dept")["salary"].mean())
# 多指标聚合print(df.groupby("dept").agg({ "salary": ["mean", "max", "count"], "age": "mean",}))# 输出:# salary age# mean max count mean# dept# Eng 70000.0 80000 2 32.5# HR 52500.0 55000 2 26.5# 缺失值处理df2 = pd.DataFrame({"a": [1, None, 3], "b": [4, 5, None]})print(df2.isnull().sum()) # 统计每列缺失值数量df2 = df2.fillna(0) # 用 0 填充缺失值
# 去重df = df.drop_duplicates()
# 类型转换df["age"] = df["age"].astype(float)
# 重命名列df = df.rename(columns={"dept": "department"})Matplotlib:可视化
Section titled “Matplotlib:可视化”数据可视化是理解数据最直接的方式。
import matplotlib.pyplot as plt
# 折线图x = np.linspace(0, 10, 100)plt.plot(x, np.sin(x), label="sin(x)")plt.plot(x, np.cos(x), label="cos(x)")plt.xlabel("x")plt.ylabel("y")plt.title("sin and cos")plt.legend()plt.grid(True, alpha=0.3)plt.show()
# 散点图x = np.random.randn(100)y = 2 * x + np.random.randn(100) * 0.5plt.scatter(x, y, alpha=0.6, c=np.abs(x), cmap="viridis")plt.colorbar(label="|x|")plt.show()
# 直方图data = np.random.randn(1000)plt.hist(data, bins=30, edgecolor="white", alpha=0.7)plt.axvline(data.mean(), color="red", linestyle="--", label=f"均值={data.mean():.2f}")plt.legend()plt.show()实战:鸢尾花数据集分析
Section titled “实战:鸢尾花数据集分析”把三个库串起来,完成一个完整的分析流程:
flowchart LR A[加载数据] --> B[DataFrame 整理] B --> C[分组统计] B --> D[可视化探索] C --> E[得出结论] D --> Efrom sklearn.datasets import load_irisimport pandas as pdimport matplotlib.pyplot as plt
# 1. 加载数据iris = load_iris()df = pd.DataFrame(iris.data, columns=iris.feature_names)df["species"] = iris.target
# 2. 数据概览:每个物种各特征的平均值print(df.groupby("species").mean())
# 3. 可视化:不同物种的特征分布fig, axes = plt.subplots(1, 3, figsize=(12, 4))for i, feature in enumerate(iris.feature_names[:3]): for species in range(3): subset = df[df["species"] == species] axes[i].hist(subset[feature], alpha=0.5, label=iris.target_names[species], bins=15) axes[i].set_title(feature) axes[i].legend(fontsize=7)plt.tight_layout()plt.show()NumPy → 数值计算、矩阵运算(底层)Pandas → 数据整理、统计分析(中层)Matplotlib → 数据可视化(上层)三个库配合使用,构成了 Python 数据分析的完整工具链。