当前位置：首页 > news >正文

esp32-s3训练自己的数据进行目标检测、图像分类

news 2025/7/3 11:38:24

esp32-s3训练自己的数据进行目标检测、图像分类

- 一、下载项目
- 二、环境
- 三、训练和导出模型
- 四、部署模型
- 五、存在的问题

esp-idf的安装参考我前面的文章： esp32cam和esp32-s3烧录human_face_detect实现人脸识别

一、下载项目

训练、转换模型：ModelAssistant(main)
部署模型：sscma-example-esp32(1.0.0)
说明文档：sscma-model-zoo

二、环境

python3.8 + CUDA11.7 + esp-idf5.0
# 主要按照ModelAssistant/requirements_cuda.txt，如果训练时有库不兼容的问题可参考下方
torch                        2.0.0+cu117
torchaudio                   2.0.1+cu117
torchvision                  0.15.1+cu117
yapf                         0.40.2
typing_extensions            4.5.0
tensorboard                  2.13.0
tensorboard-data-server      0.7.2
tensorflow                   2.13.0
keras                        2.13.1
tensorflow-estimator         2.13.0
tensorflow-intel             2.13.0
tensorflow-io-gcs-filesystem 0.31.0
sscma                        2.0.0rc3
setuptools                   60.2.0
rich                         13.4.2
Pillow                       9.4.0
mmcls                        1.0.0rc6
mmcv                         2.0.0
mmdet                        3.0.0
mmengine                     0.10.1
mmpose                       1.2.0
mmyolo                       0.5.0

三、训练和导出模型

step 1: 将voc格式的标注文件转换为edgelab的训练格式，并按8：2的比例划分为训练集和验证集

import os
import json
import pandas as pd
from xml.etree import ElementTree as ET
from PIL import Image
import shutil
import random
from tqdm import tqdm# Set paths
voc_path = 'F:/datasets/VOCdevkit/VOC2007'
train_path = 'F:/edgelab/ModelAssistant/datasets/myself/train'
valid_path = 'F:/edgelab/ModelAssistant/datasets/meself/valid'# 只读取有目标的，且属于需要训练的类别
classes = ["face"]# Create directories if not exist
if not os.path.exists(train_path):os.makedirs(train_path)
if not os.path.exists(valid_path):os.makedirs(valid_path)# Get list of image files
image_files = os.listdir(os.path.join(voc_path, 'JPEGImages'))
random.seed(0)
random.shuffle(image_files)# Split data into train and valid
train_files = image_files[:int(len(image_files)*0.8)]
valid_files = image_files[int(len(image_files)*0.8):]# Convert train data to COCO format
train_data = {'categories': [], 'images': [], 'annotations': []}
train_ann_id = 0
train_cat_id = 0
img_id = 0
train_categories = {}
for file in tqdm(train_files):# Add annotationsxml_file = os.path.join(voc_path, 'Annotations', file[:-4] + '.xml')tree = ET.parse(xml_file)root = tree.getroot()for obj in root.findall('object'):category = obj.find('name').textif category not in classes:continueif category not in train_categories:train_categories[category] = train_cat_idtrain_cat_id += 1category_id = train_categories[category]bbox = obj.find('bndbox')x1 = int(bbox.find('xmin').text)y1 = int(bbox.find('ymin').text)x2 = int(bbox.find('xmax').text)y2 = int(bbox.find('ymax').text)width = x2 - x1height = y2 - y1ann_info = {'id': train_ann_id, 'image_id': img_id, 'category_id': category_id, 'bbox': [x1, y1, width, height],'area': width*height, 'iscrowd': 0}train_data['annotations'].append(ann_info)train_ann_id += 1if len(root.findall('object')):# 只有有目标的图片才加进来image_id = img_idimg_id += 1image_file = os.path.join(voc_path, 'JPEGImages', file)shutil.copy(image_file, os.path.join(train_path, file))img = Image.open(image_file)image_info = {'id': image_id, 'file_name': file, 'width': img.size[0], 'height': img.size[1]}train_data['images'].append(image_info)# Add categories
for category, category_id in train_categories.items():train_data['categories'].append({'id': category_id, 'name': category})# Save train data to file
with open(os.path.join(train_path, '_annotations.coco.json'), 'w') as f:json.dump(train_data, f, indent=4)# Convert valid data to COCO format
valid_data = {'categories': [], 'images': [], 'annotations': []}
valid_ann_id = 0
img_id = 0
for file in tqdm(valid_files):# Add annotationsxml_file = os.path.join(voc_path, 'Annotations', file[:-4] + '.xml')tree = ET.parse(xml_file)root = tree.getroot()for obj in root.findall('object'):category = obj.find('name').textif category not in classes:continuecategory_id = train_categories[category]bbox = obj.find('bndbox')x1 = int(bbox.find('xmin').text)y1 = int(bbox.find('ymin').text)x2 = int(bbox.find('xmax').text)y2 = int(bbox.find('ymax').text)width = x2 - x1height = y2 - y1ann_info = {'id': valid_ann_id, 'image_id': img_id, 'category_id': category_id, 'bbox': [x1, y1, width, height],'area': width*height, 'iscrowd': 0}valid_data['annotations'].append(ann_info)valid_ann_id += 1if len(root.findall('object')):# Add imageimage_id = img_idimg_id += 1image_file = os.path.join(voc_path, 'JPEGImages', file)shutil.copy(image_file, os.path.join(valid_path, file))img = Image.open(image_file)image_info = {'id': image_id, 'file_name': file, 'width': img.size[0], 'height': img.size[1]}valid_data['images'].append(image_info)# Add categories
valid_data['categories'] = train_data['categories']# Save valid data to file
with open(os.path.join(valid_path, '_annotations.coco.json'), 'w') as f:json.dump(valid_data, f, indent=4)

step 2: 参考Face Detection - Swift-YOLO下载模型权重文件和训练

python tools/train.py configs/yolov5/yolov5_tiny_1xb16_300e_coco.py \
--cfg-options  \work_dir=work_dirs/face_96 \num_classes=3 \epochs=300  \height=96 \width=96 \batch=128 \data_root=datasets/face/ \load_from=datasets/face/pretrain.pth

step 3: 训练过程可视化tensorboard

cd work_dirs/face_96/20231219_181418/vis_data
tensorboard --logdir=./

然后按照提示打开http://localhost:6006/
在这里插入图片描述

step 4: 导出模型

python tools/export.py configs/yolov5/yolov5_tiny_1xb16_300e_coco.py ./work_dirs/face_96/best_coco_bbox_mAP_epoch_300.pth --target tflite onnx
--cfg-options  \work_dir=work_dirs/face_96 \num_classes=3 \epochs=300  \height=96 \width=96 \batch=128 \data_root=datasets/face/ \load_from=datasets/face/pretrain.pth

这样就会在./work_dirs/face_96路径下生成best_coco_bbox_mAP_epoch_300_int8.tflite文件了。

四、部署模型

step 1: 将best_coco_bbox_mAP_epoch_300_int8.tflite复制到F:\edgelab\sscma-example-esp32-1.0.0\model_zoo路径下
step 2: 参照edgelab-example-esp32-训练和部署一个FOMO模型将模型转换为C语言文件，并将其放入到F:\edgelab\sscma-example-esp32-1.0.0\components\modules\model路径下

python tools/tflite2c.py --input ./model_zoo/best_coco_bbox_mAP_epoch_300_int8.tflite --name yolo --output_dir ./components/modules/model --classes face

这样会生成./components/modules/model/yolo_model_data.cpp和yolo_model_data.h两个文件。

step 3: 利用idf烧录程序

fb_gfx_printf(frame, yolo.x - yolo.w / 2, yolo.y - yolo.h/2 - 5, 0x1FE0, "%s:%d", g_yolo_model_classes[yolo.target], yolo.confidence);

打开esp-idf cmd

cd F:\edgelab\sscma-example-esp32-1.0.0\examples\yolo
idf.py set-target esp32s3
idf.py menuconfig

在这里插入图片描述勾选上方的这个选项不然报错

E:/Softwares/Espressif/frameworks/esp-idf-v5.0.4/components/driver/deprecated/driver/i2s.h:27:2: warning: #warning "This set of I2S APIs has been deprecated, please include 'driver/i2s_std.h', 'driver/i2s_pdm.h' or 'driver/i2s_tdm.h' instead. if you want to keep using the old APIs and ignore this warning, you can enable 'Suppress leagcy driver deprecated warning' option under 'I2S Configuration' menu in Kconfig" [-Wcpp]27 | #warning "This set of I2S APIs has been deprecated, \|  ^~~~~~~
ninja: build stopped: subcommand failed.
ninja failed with exit code 1, output of the command is in the F:\edgelab\sscma-example-esp32-1.0.0\examples\yolo\build\log\idf_py_stderr_output_27512 and F:\edgelab\sscma-example-esp32-1.0.0\examples\yolo\build\log\idf_py_stdout_output_27512

idf.py flash monitor -p COM3

在这里插入图片描述
lcd端也能实时显示识别结果，输入大小为96x96时推理时间大概200ms，192x192时时间大概660ms

五、存在的问题

该链路中量化是比较简单的，在我的数据集上量化后精度大打折扣，应该需要修改量化算法，后续再说吧。

量化前
量化后

esp32-s3训练自己的数据进行目标检测、图像分类

esp32-s3训练自己的数据进行目标检测、图像分类一、下载项目二、环境三、训练和导出模型四、部署模型五、存在的问题 esp-idf的安装参考我前面的文章： esp32cam和esp32-s3烧录human_face_detect实现人脸识别一、下载项目训练、转换模型：ModelAssist…...

编程日记 2023/12/22 14:35:18

华为设备VRP基础

交换机可以隔离冲突域，路由器可以隔离广播域，这两种设备在企业网络中应用越来越广泛。随着越来越多的终端接入到网络中，网络设备的负担也越来越重，这时网络设备可以通过华为专有的VRP系统来提升运行效率。通用路由平台VRP&#xf…...

编程日记 2023/12/22 14:33:16

论文笔记 | ICLR 2023 WikiWhy：回答和解释因果问题

文章目录一、前言二、主要内容三、总结🍉 CSDN 叶庭云：https://yetingyun.blog.csdn.net/ 一、前言 ICLR 2023 | Accept: notable-top-5%：《WikiWhy: Answering and Explaining Cause-and-Effect Questions》一段话总结：WikiWhy 是一个新的 QA 数据集，围绕一个新的任务…...

编程日记 2023/12/22 14:31:14

LC24. 两两交换链表中的节点

代码随想录 class Solution {// 举例子:假设两个节点 1 -> 2// 那么 head 1; next 2; next.next null// 那么swapPairs(next.next),传入的是null,再下一次递归中直接返回null// 因此 newNode null// 所以 next.next head; > 2.next 1; 2 -> 1// head.next…...

编程日记 2023/12/22 14:30:13

使用redis-rds-tools 工具分析redis rds文件

redis-rdb-tools安装部署及使用发布时间：2020-07-28 12:33:12 阅读：29442 作者：苏黎世1995 栏目：关系型数据库活动：开发者测试专用服务器限时活动，0元免费领，库存有限，领完即止&…...

编程日记 2023/12/22 14:26:09

C# Onnx yolov8 plane detection

C# Onnx yolov8 plane detection 效果模型信息 Model Properties ------------------------- date：2023-12-22T10:57:49.823820 author：Ultralytics task：detect license：AGPL-3.0 https://ultralytics.com/license version&am…...

编程日记 2023/12/22 14:21:04

Oracle定时任务的创建与禁用/删除

在开始操作之前，先从三W开始，即我常说的what 是什么；why 为什么使用；how 如何使用。一、Oracle定时器是什么 Oracle定时器是一种用于在特定时间执行任务或存储过程的工具，可以根据需求设置不同的时间段和频率来执行…...

编程日记 2023/12/22 14:17:59

Asp.Net Core 项目中常见中间件调用顺序

常用的 AspNetCore 项目中间件有这些，调用顺序如下图所示： 最后的 Endpoint 就是最终生成响应的中间件。 Configure调用如下： public void Configure(IApplicationBuilder app, IWebHostEnvironment env){if (env.IsDevelopment()){app.UseD…...

编程日记 2023/12/22 14:12:54

【JVM】一、认识JVM

文章目录 1、虚拟机2、Java虚拟机3、JVM的整体结构4、Java代码的执行流程5、JVM的分类6、JVM的生命周期 1、虚拟机虚拟机，Virtual Machine，一台虚拟的计算机，用来执行虚拟计算机指令。分为： 系统虚拟机：如VMware&am…...

编程日记 2023/12/22 14:11:53

[SWPUCTF 2021 新生赛]Do_you_know_http已

打开环境它说用WLLM浏览器打开，使用BP抓包，发送到重发器修改User-Agent 下一步，访问a.php 这儿他说添加一个本地地址，它给了一个183.224.40.160，我用了发现没用，然后重新添加一个地址：X-Forwa…...

编程日记 2023/12/22 14:10:52

hadoop01_完全分布式搭建

hadoop完全分布式搭建 1 完全分布式介绍 Hadoop运行模式包括：本地模式（计算的数据存在Linux本地，在一台服务器上自己测试）、伪分布式模式（和集群接轨 HDFS yarn，在一台服务器上执行）、完全分…...

编程日记 2023/12/22 14:00:42

【每日一题】得到山形数组的最少删除次数

文章目录 Tag题目来源解题思路方法一：最长递增子序列写在最后 Tag 【最长递增子序列】【数组】【2023-12-22】题目来源 1671. 得到山形数组的最少删除次数解题思路方法一：最长递增子序列前后缀分解根据前后缀思想，以 nums[i] 为山…...

编程日记 2023/12/22 13:57:39

2023年，为什么汽车依然有很多小毛病？

汽车出现小毛病是一个复杂的问题，其原因涉及到汽车本身的设计、制造质量、维护保养以及使用环境等多个方面。只有汽车制造商、车主和社会各界共同努力，才能够减少汽车的小毛病，提高汽车的可靠性和安全性。比如，汽车的维护和保养…...

编程日记 2023/12/22 13:53:36

yocto系列讲解[实战篇]93 - 添加Qtwebengine和Browser实例

By: fulinux E-mail: fulinux@sina.com Blog: https://blog.csdn.net/fulinus 喜欢的盆友欢迎点赞和订阅！你的喜欢就是我写作的动力！目录概述集成meta-qt5移植过程中的问题问题1：virtual/libgl set to mesa, not mesa-gl问题2：dmabuf-server-buffer tries to use undecl…...

编程日记 2023/12/22 13:50:33

Python实验报告十一、自定义类模拟三维向量及其运算

一、实验目的： 1、了解如何定义一个类。 2、了解如何定义类的私有数据成员和成员方法。 3、了解如何使用自定义类实例化对象。二、实验内容： 定义一个三维向量类，并定义相应的特殊方法实现两个该类对象之间的加、减运算（要…...

编程日记 2023/12/22 13:49:32

机器学习 | 聚类Clustering 算法

物以类聚人以群分。什么是聚类呢？ 1、核心思想和原理聚类的目的同簇高相似度不同簇高相异度同类尽量相聚不同类尽量分离聚类和分类的区别分类 classification 监督学习训练获得分类器预测未知数据聚类 clustering 无监督学习，不关心类别标签 …...

编程日记 2023/12/22 13:48:31

IntelliJ IDEA 2023.3 新功能介绍

IntelliJ IDEA 2023.3 在众多领域进行了全面的改进，引入了许多令人期待的功能和增强体验。以下是该版本的一些关键亮点： IntelliJ IDEA mac版下载 macappbox.com/a/intellij-idea-for-mac.html 1. AI Assistant 的全面推出 IntelliJ IDEA 2023.3 中&am…...

编程日记 2023/12/22 13:46:30

2. 行为模式 - 命令模式

亦称： 动作、事务、Action、Transaction、Command 意图命令模式是一种行为设计模式， 它可将请求转换为一个包含与请求相关的所有信息的独立对象。该转换让你能根据不同的请求将方法参数化、延迟请求执行或将其放入队列中， 且能实现可撤销…...

编程日记 2023/12/22 13:45:29

Java智慧工地源码 SAAS智慧工地源码智慧工地管理可视化平台源码带移动APP

一、系统主要功能介绍系统功能介绍： 【项目人员管理】 1. 项目管理：项目名称、施工单位名称、项目地址、项目地址、总造价、总面积、施工准可证、开工日期、计划竣工日期、项目状态等。 2. 人员信息管理：支持身份证及人脸信息采集&#…...

编程日记 2023/12/22 13:44:28

php学习02-php标记风格

<?php echo "这是xml格式风格" ?><script language"php">echo 脚本风格标记 </script><% echo "这是asp格式风格" %>推荐使用xml格式风格如果要使用简短风格和ASP风格，需要在php.ini中对其进行配置&#…...

编程日记 2023/12/22 13:43:27

DockerHub与私有镜像仓库在容器化中的应用与管理

哈喽，大家好，我是左手python！ Docker Hub的应用与管理 Docker Hub的基本概念与使用方法 Docker Hub是Docker官方提供的一个公共镜像仓库，用户可以在其中找到各种操作系统、软件和应用的镜像。开发者可以通过Docker Hub轻松获取所…...

编程新知 2025/6/27 0:59:29

使用 SymPy 进行向量和矩阵的高级操作

在科学计算和工程领域，向量和矩阵操作是解决问题的核心技能之一。Python 的 SymPy 库提供了强大的符号计算功能，能够高效地处理向量和矩阵的各种操作。本文将深入探讨如何使用 SymPy 进行向量和矩阵的创建、合并以及维度拓展等操作，并通过具体…...

编程新知 2025/7/3 3:01:17

【Go语言基础【12】】指针：声明、取地址、解引用

文章目录零、概述：指针 vs. 引用（类比其他语言）一、指针基础概念二、指针声明与初始化三、指针操作符1. &：取地址（拿到内存地址）2. *：解引用（拿到值） 四、空指针&am…...

编程新知 2025/6/21 2:18:57

MySQL JOIN 表过多的优化思路

当 MySQL 查询涉及大量表 JOIN 时，性能会显著下降。以下是优化思路和简易实现方法： 一、核心优化思路减少 JOIN 数量数据冗余：添加必要的冗余字段（如订单表直接存储用户名）合并表：将频繁关联的小表合并成…...

编程新知 2025/6/16 23:36:31

宇树科技，改名了！

提到国内具身智能和机器人领域的代表企业，那宇树科技（Unitree）必须名列其榜。最近，宇树科技的一项新变动消息在业界引发了不少关注和讨论，即： 宇树向其合作伙伴发布了一封公司名称变更函称，因…...

编程新知 2025/6/25 1:25:32

[大语言模型]在个人电脑上部署ollama 并进行管理,最后配置AI程序开发助手.

ollama官网: 下载 https://ollama.com/ 安装查看可以使用的模型 https://ollama.com/search 例如 https://ollama.com/library/deepseek-r1/tags # deepseek-r1:7bollama pull deepseek-r1:7b改token数量为409622 16384 ollama命令说明 ollama serve #&#xff1a…...

编程新知 2025/6/27 0:08:08

WebRTC从入门到实践 - 零基础教程

WebRTC从入门到实践 - 零基础教程目录 WebRTC简介基础概念工作原理开发环境搭建基础实践三个实战案例常见问题解答 1. WebRTC简介 1.1 什么是WebRTC？ WebRTC（Web Real-Time Communication）是一个支持网页浏览器进行实时语音…...

编程新知 2025/6/21 6:26:12

Linux 下 DMA 内存映射浅析

序系统 I/O 设备驱动程序通常调用其特定子系统的接口为 DMA 分配内存，但最终会调到 DMA 子系统的dma_alloc_coherent()/dma_alloc_attrs() 等接口。关于 dma_alloc_coherent 接口详细的代码讲解、调用流程，可以参考这篇文章，我觉得写的非常…...

编程新知 2025/7/2 18:36:35

算法—栈系列

一：删除字符串中的所有相邻重复项 class Solution { public:string removeDuplicates(string s) {stack<char> st;for(int i 0; i < s.size(); i){char target s[i];if(!st.empty() && target st.top())st.pop();elsest.push(s[i]);}string ret…...

编程新知 2025/7/3 5:29:11

Springboot 高校报修与互助平台小程序

一、前言随着我国经济迅速发展，人们对手机的需求越来越大，各种手机软件也都在被广泛应用，但是对于手机进行数据信息管理，对于手机的各种软件也是备受用户的喜爱，高校报修与互助平台小程序被用户普遍使用，为…...

编程新知 2025/7/2 0:46:41

esp32-s3训练自己的数据进行目标检测、图像分类

一、下载项目

二、环境

三、训练和导出模型

四、部署模型

五、存在的问题

相关文章：