当前位置：首页 > news >正文

golang map真有那么随机吗？——map遍历研究

news 2025/7/8 16:13:25

在随机选取map中元素时，本想用map遍历的方式来返回，但是却并没有通过测试。

那么难道map的遍历并不是那么的随机吗？

以下代码参考go1.18

hiter是map遍历的结构，主要记录了当前遍历的元素、开始位置等来完成整个遍历过程

// A hash iteration structure.
// If you modify hiter, also change cmd/compile/internal/reflectdata/reflect.go
// and reflect/value.go to match the layout of this structure.
type hiter struct {// 指向下一个遍历key的地址key         unsafe.Pointer // Must be in first position.  Write nil to indicate iteration end (see cmd/compile/internal/walk/range.go).// 指向下一个遍历value的地址elem        unsafe.Pointer // Must be in second position (see cmd/compile/internal/walk/range.go).// map类型t           *maptype// map headerh           *hmap// 初始化时指向的bucketbuckets     unsafe.Pointer // bucket ptr at hash_iter initialization time// 当前遍历到的bmapbptr        *bmap          // current bucketoverflow    *[]*bmap       // keeps overflow buckets of hmap.buckets aliveoldoverflow *[]*bmap       // keeps overflow buckets of hmap.oldbuckets alive// 开始桶startBucket uintptr        // bucket iteration started at// 桶内偏移量offset      uint8          // intra-bucket offset to start from during iteration (should be big enough to hold bucketCnt-1)// 是否从头遍历了wrapped     bool           // already wrapped around from end of bucket array to beginningB           uint8// 正在遍历的槽位i           uint8// 正在遍历的桶位bucket      uintptr// 用于扩容时进行检查checkBucket uintptr
}

mapiterinit为开始遍历的方法，主要是确定初始遍历的位置

// mapiterinit initializes the hiter struct used for ranging over maps.
// The hiter struct pointed to by 'it' is allocated on the stack
// by the compilers order pass or on the heap by reflect_mapiterinit.
// Both need to have zeroed hiter since the struct contains pointers.
func mapiterinit(t *maptype, h *hmap, it *hiter) {// 若map为空，则跳过遍历过程it.t = tif h == nil || h.count == 0 {return}if unsafe.Sizeof(hiter{})/goarch.PtrSize != 12 {throw("hash_iter size incorrect") // see cmd/compile/internal/reflectdata/reflect.go}it.h = h// grab snapshot of bucket state// 迭代器快照记录map桶信息it.B = h.Bit.buckets = h.bucketsif t.bucket.ptrdata == 0 {// Allocate the current slice and remember pointers to both current and old.// This preserves all relevant overflow buckets alive even if// the table grows and/or overflow buckets are added to the table// while we are iterating.h.createOverflow()it.overflow = h.extra.overflowit.oldoverflow = h.extra.oldoverflow}// decide where to start// 开始bucket选择随机数的低B位// 偏移量选择随机数高B位与桶数量，显然这个桶数量是不包括溢出桶的r := uintptr(fastrand())if h.B > 31-bucketCntBits {r += uintptr(fastrand()) << 31}it.startBucket = r & bucketMask(h.B)it.offset = uint8(r >> h.B & (bucketCnt - 1))// iterator state// 更新迭代器桶为初始桶it.bucket = it.startBucket// Remember we have an iterator.// Can run concurrently with another mapiterinit().// 标记可能有迭代正在使用桶和旧桶if old := h.flags; old&(iterator|oldIterator) != iterator|oldIterator {atomic.Or8(&h.flags, iterator|oldIterator)}mapiternext(it)
}

从上面的代码分析我们便可以看出随机选取的元素并不是真的随机，溢出桶并不包含在随机选择的范围里面

在具体的遍历过程，存在以下疑问

如果在扩容中，如何进行遍历？
如何保证不遗漏？
如何防止重复遍历？

func mapiternext(it *hiter) {h := it.h// 如果标记已经写入，则抛出并发迭代写入错误if h.flags&hashWriting != 0 {throw("concurrent map iteration and map write")}t := it.tbucket := it.bucketb := it.bptri := it.icheckBucket := it.checkBucketnext:if b == nil {// 如果再次遇到开始bucket且是从头遍历的，则说明迭代结束，返回if bucket == it.startBucket && it.wrapped {// end of iterationit.key = nilit.elem = nilreturn}// 如果正在迁移过程中，且老桶没被迁移，采用老桶if h.growing() && it.B == h.B {// Iterator was started in the middle of a grow, and the grow isn't done yet.// If the bucket we're looking at hasn't been filled in yet (i.e. the old// bucket hasn't been evacuated) then we need to iterate through the old// bucket and only return the ones that will be migrated to this bucket.oldbucket := bucket & it.h.oldbucketmask()b = (*bmap)(add(h.oldbuckets, oldbucket*uintptr(t.bucketsize)))// bucket未迁移，记录bucket// checkBucket在当前map处于迁移而bucket未迁移时，为当前bucket// 否则为noCheckif !evacuated(b) {checkBucket = bucket} else {b = (*bmap)(add(it.buckets, bucket*uintptr(t.bucketsize)))checkBucket = noCheck}} else {// map处于未迁移，或者bucket迁移完成，采用新桶b = (*bmap)(add(it.buckets, bucket*uintptr(t.bucketsize)))checkBucket = noCheck}// 推进到下一桶bucket++// 遍历到最后一个桶，要绕回0桶继续遍历if bucket == bucketShift(it.B) {bucket = 0it.wrapped = true}i = 0}// 遍历桶内元素for ; i < bucketCnt; i++ {// 从offset槽开始offi := (i + it.offset) & (bucketCnt - 1)// 跳过空槽if isEmpty(b.tophash[offi]) || b.tophash[offi] == evacuatedEmpty {// TODO: emptyRest is hard to use here, as we start iterating// in the middle of a bucket. It's feasible, just tricky.continue}// 获取元素key、valuek := add(unsafe.Pointer(b), dataOffset+uintptr(offi)*uintptr(t.keysize))if t.indirectkey() {k = *((*unsafe.Pointer)(k))}e := add(unsafe.Pointer(b), dataOffset+bucketCnt*uintptr(t.keysize)+uintptr(offi)*uintptr(t.elemsize))// 扩容迁移时过滤掉不属于当前指向新桶的旧桶元素if checkBucket != noCheck && !h.sameSizeGrow() {// Special case: iterator was started during a grow to a larger size// and the grow is not done yet. We're working on a bucket whose// oldbucket has not been evacuated yet. Or at least, it wasn't// evacuated when we started the bucket. So we're iterating// through the oldbucket, skipping any keys that will go// to the other new bucket (each oldbucket expands to two// buckets during a grow).// 若key是有效的if t.reflexivekey() || t.key.equal(k, k) {// If the item in the oldbucket is not destined for// the current new bucket in the iteration, skip it.// 如果旧桶中的项在迭代中不打算用于当前的新桶，则跳过它。hash := t.hasher(k, uintptr(h.hash0))if hash&bucketMask(it.B) != checkBucket {continue}} else {// 对k！=k，也就是nil之类的，判断是否属于该新桶// 不是，则跳过// Hash isn't repeatable if k != k (NaNs).  We need a// repeatable and randomish choice of which direction// to send NaNs during evacuation. We'll use the low// bit of tophash to decide which way NaNs go.// NOTE: this case is why we need two evacuate tophash// values, evacuatedX and evacuatedY, that differ in// their low bit.if checkBucket>>(it.B-1) != uintptr(b.tophash[offi]&1) {continue}}}// 如果当前桶未扩容迁移，或者是每次hash不一致的key，获取到key、value添加到迭代器中if (b.tophash[offi] != evacuatedX && b.tophash[offi] != evacuatedY) ||!(t.reflexivekey() || t.key.equal(k, k)) {// This is the golden data, we can return it.// OR// key!=key, so the entry can't be deleted or updated, so we can just return it.// That's lucky for us because when key!=key we can't look it up successfully.it.key = kif t.indirectelem() {e = *((*unsafe.Pointer)(e))}it.elem = e} else {// 数据已经迁移情况下，处理键已被删除、更新或删除并重新插入的情况，定位数据，最后添加遍历key、value// The hash table has grown since the iterator was started.// The golden data for this key is now somewhere else.// Check the current hash table for the data.// This code handles the case where the key// has been deleted, updated, or deleted and reinserted.// NOTE: we need to regrab the key as it has potentially been// updated to an equal() but not identical key (e.g. +0.0 vs -0.0).rk, re := mapaccessK(t, h, k)if rk == nil {continue // key has been deleted}it.key = rkit.elem = re}// 迭代器记录进度it.bucket = bucketif it.bptr != b { // avoid unnecessary write barrier; see issue 14921it.bptr = b}it.i = i + 1it.checkBucket = checkBucketreturn}// 遍历溢出桶b = b.overflow(t)i = 0goto next
}

通过以上代码分析，可以看出：

在扩容时遍历，
- 如果当前遍历的桶已经迁移好了，那么取新桶
- 如果仍然处于旧桶，则取旧桶。
  
  但值得注意的是要过滤掉那些不属于该新桶的旧桶元素。因为旧桶在扩容迁移时会分为两块，当前指向的新桶只属于其中之一
bucket从初始桶逐渐递增，保证正常桶都能遍历到。此外也保证了完整遍历溢出桶，直到溢出桶为空
通过记录是否从头遍历的标志和起始bucket，以及在扩容过程中过滤不属于该新桶的元素来保证不会重复遍历

Ref

https://zhuanlan.zhihu.com/p/597348765
https://www.cnblogs.com/cnblogs-wangzhipeng/p/13292524.html
https://qcrao.com/post/dive-into-go-map/

golang map真有那么随机吗？——map遍历研究

在随机选取map中元素时，本想用map遍历的方式来返回，但是却并没有通过测试。那么难道map的遍历并不是那么的随机吗？ 以下代码参考go1.18 hiter是map遍历的结构，主要记录了当前遍历的元素、开始位置等来完成整个遍历过程 // A ha…...

编程日记 2024/1/27 1:13:15

详细分析对比copliot和ChatGPT的差异

Copilot 和 ChatGPT 是两种不同的AI工具，分别在不同领域展现出了强大的功能和潜力： GitHub Copilot 定位与用途：GitHub Copilot 是由GitHub（现为微软子公司）和OpenAI合作开发的一款智能代码辅助工具。它主要集成于Visu…...

编程日记 2024/1/27 1:10:13

TENT:熵最小化的Fully Test-Time Adaption

摘要在测试期间，模型必须自我调整以适应新的和不同的数据。在这种完全自适应测试时间的设置中，模型只有测试数据和它自己的参数。我们建议通过test entropy minimization (tent[1])来适应:我们通过其预测的熵来优化模型的置信度。我们的方法估计归一化…...

编程日记 2024/1/27 1:06:09

研发日记，Matlab/Simulink避坑指南(五)——CAN解包 DLC Bug

文章目录前言背景介绍问题描述分析排查解决方案总结前言见《研发日记，Matlab/Simulink避坑指南（一）——Data Store Memory模块执行时序Bug》见《研发日记，Matlab/Simulink避坑指南(二)——非对称数据溢出Bug》见《…...

编程日记 2024/1/27 1:02:05

机器人3D视觉引导半导体塑封上下料

半导体塑封上下料是封装工艺中的重要环节，直接影响到产品的质量和性能。而3D视觉引导技术的引入，使得这一过程更加高效、精准。它不仅提升了生产效率，减少了人工操作的误差，还为半导体封装技术的智能化升级奠定了坚实的基础。传统…...

编程日记 2024/1/27 1:01:04

（十二）Head first design patterns代理模式（c++）

代理模式代理模式：创建一个proxy对象，并为这个对象提供替身或者占位符以对这个对象进行控制。典型例子：智能指针... 例子：比如说有一个talk接口，所有的people需要实现talk接口。但有些人有唱歌技能。不能在talk接…...

编程日记 2024/1/27 0:57:00

C++从零开始的打怪升级之路(day21)

这是关于一个普通双非本科大一学生的C的学习记录贴在此前，我学了一点点C语言还有简单的数据结构，如果有小伙伴想和我一起学习的，可以私信我交流分享学习资料那么开启正题今天分享的是关于vector的题目 1.删除有序数组中的重复项 26. …...

编程日记 2024/1/27 0:55:59

《设计模式的艺术》笔记 - 观察者模式

介绍观察者模式定义对象之间的一种一对多依赖关系，使得每当一个对象状态发生改变时，其相关依赖对象皆得到通知并被自动更新。实现 myclass.h // // Created by yuwp on 2024/1/12. //#ifndef DESIGNPATTERNS_MYCLASS_H #define DESIGNPATTERNS_MYCLA…...

编程日记 2024/1/27 0:50:55

Java如何对OSS存储引擎的Bucket进行创建【OSS学习】

在前面学会了如何开通OSS，对OSS的一些基本操作，接下来记录一下如何通过Java代码通过SDK对OSS存储引擎里面的Bucket存储空间进行创建。目录 1、先看看OSS： 2、代码编写： 3、运行效果： 1、先看看OSS： 此…...

编程日记 2024/1/27 0:49:55

ModuleNotFoundError: No module named ‘half_json‘

问题: ModuleNotFoundError: No module named ‘half_json’ 原因: 缺少jsonfixer包解决方法: pip install jsonfixerjson修正包地址: https://github.com/half-pie/half-json...

编程日记 2024/1/27 0:47:52

深入探究 Android 内存泄漏检测原理及 LeakCanary 源码分析

深入探究 Android 内存泄漏检测原理及 LeakCanary 源码分析一、什么是内存泄漏二、内存泄漏的常见原因三、我为什么要使用 LeakCanary四、LeakCanary介绍五、LeakCanary 的源码分析及其核心代码六、LeakCanary 使用示例一、什么是内存泄漏在基于 Java 的运行时中&#xff0…...

编程日记 2024/1/27 0:43:49

Linux CentOS使用Docker搭建laravel项目环境（实践案例详细说明）

一、安装docker # 1、更新系统软件包： sudo yum update# 2、安装Docker依赖包 sudo yum install -y yum-utils device-mapper-persistent-data lvm2# 3、添加Docker的yum源： sudo yum-config-manager --add-repo https://download.docker.com/linux/cen…...

编程日记 2024/1/27 0:42:48

第六课：Prompt

文章目录第六课：Prompt1、学习总结：Prompt介绍预训练和微调模型回顾挑战 Pre-train, Prompt, PredictPrompting是什么?prompting流程prompt设计课程ppt及代码地址 2、学习心得：3、经验分享：4、课程反馈：5、使用Mind…...

编程日记 2024/1/27 0:40:46

网络安全（初版，以后会不断更新）

1.网络安全常识及术语资产任何对组织业务具有价值的信息资产，包括计算机硬件、通信设施、IT 环境、数据库、软件、文档资料、信息服务和人员等。漏洞上边提到的“永恒之蓝”就是windows系统的漏洞漏洞又被称为脆弱性或弱点（Weakness）&a…...

编程日记 2024/1/27 0:37:44

开始学习Vue2（脚手架，组件化开发）

一、单页面应用程序单页面应用程序（英文名：Single Page Application）简称 SPA，顾名思义，指的是一个 Web 网站中只有唯一的一个 HTML 页面，所有的功能与交互都在这唯一的一个页面内完成。二、vue-cli …...

编程日记 2024/1/27 0:34:40

平替heygen的开源音频克隆工具—OpenVoice

截止2024-1-26日，全球范围内语音唇形实现最佳的应该算是heygen，可惜不但要魔法，还需要银子；那么有没有可以平替的方案，答案是肯定的。方案1： 采用国内星火大模型训练自己的声音，然后再用下面…...

编程日记 2024/1/27 0:31:37

【自动化测试】读写64位操作系统的注册表

自动化测试经常需要修改注册表很多系统的设置（比如：IE的设置）都是存在注册表中。桌面应用程序的设置也是存在注册表中。所以做自动化测试的时候，经常需要去修改注册表 Windows注册表简介注册表编辑器在 C:\Windows\regedit…...

编程日记 2024/1/27 0:30:35

php二次开发股票系统代码：腾讯股票数据接口地址、批量获取股票信息、转换为腾讯接口指定的股票格式

1、腾讯股票数据控制器 <?php namespace app\index\controller;use think\Model; use think\Db;const BASE_URL http://aaaaaa.aaaaa.com; //腾讯数据地址class TencentStocks extends Home { //里面具体的方法 }2、请求接口返回内容 function juhecurl($url, $params f…...

编程日记 2024/1/27 0:28:33

uniapp 在static/index.html中添加全局样式

前言略在static/index.html中添加全局样式 <style>div {background-color: #ccc;} </style>static/index.html源码： <!DOCTYPE html> <html lang"zh-CN"><head><meta charset"utf-8"><meta http-…...

编程日记 2024/1/27 0:27:32

acrobat调整pdf的页码和实际页码保持一致

Acrobat版本具体操作现在拿到pdf的结构如下： pdf页码实际页码1-10页无页码数11页第1页操作，选择pdf第10页，右键点击具体设置最终效果...

编程日记 2024/1/27 0:26:32

(LeetCode 每日一题) 3442. 奇偶频次间的最大差值 I (哈希、字符串)

题目：3442. 奇偶频次间的最大差值 I 思路 ：哈希，时间复杂度0(n)。用哈希表来记录每个字符串中字符的分布情况，哈希表这里用数组即可实现。 C版本： class Solution { public:int maxDifference(string s) {int a[26]…...

编程新知 2025/7/7 5:08:22

React hook之useRef

React useRef 详解 useRef 是 React 提供的一个 Hook，用于在函数组件中创建可变的引用对象。它在 React 开发中有多种重要用途，下面我将全面详细地介绍它的特性和用法。基本概念 1. 创建 ref const refContainer useRef(initialValue);initialValu…...

编程新知 2025/6/11 15:21:26

Linux简单的操作

ls ls 查看当前目录 ll 查看详细内容 ls -a 查看所有的内容 ls --help 查看方法文档 pwd pwd 查看当前路径 cd cd 转路径 cd .. 转上一级路径 cd 名转换路径 …...

编程新知 2025/6/26 2:36:22

MMaDA: Multimodal Large Diffusion Language Models

CODE ： https://github.com/Gen-Verse/MMaDA Abstract 我们介绍了一种新型的多模态扩散基础模型MMaDA，它被设计用于在文本推理、多模态理解和文本到图像生成等不同领域实现卓越的性能。该方法的特点是三个关键创新:(i) MMaDA采用统一的扩散架构&#xf…...

编程新知 2025/7/6 2:30:53

C# SqlSugar：依赖注入与仓储模式实践

C# SqlSugar：依赖注入与仓储模式实践在 C# 的应用开发中，数据库操作是必不可少的环节。为了让数据访问层更加简洁、高效且易于维护，许多开发者会选择成熟的 ORM（对象关系映射）框架，SqlSugar 就是其中备受…...

编程新知 2025/7/5 18:24:10

爬虫基础学习day2

# 爬虫设计领域工商：企查查、天眼查短视频：抖音、快手、西瓜 ---> 飞瓜电商：京东、淘宝、聚美优品、亚马逊 ---> 分析店铺经营决策标题、排名航空：抓取所有航空公司价格 ---> 去哪儿自媒体：采集自媒体数据进…...

编程新知 2025/7/6 13:55:34

论文笔记——相干体技术在裂缝预测中的应用研究

目录相关地震知识补充地震数据的认识地震几何属性相干体算法定义基本原理第一代相干体技术：基于互相关的相干体技术（Correlation）第二代相干体技术：基于相似的相干体技术（Semblance）基于多道相似的相干体…...

编程新知 2025/7/7 16:17:48

GruntJS-前端自动化任务运行器从入门到实战

Grunt 完全指南：从入门到实战一、Grunt 是什么？ Grunt是一个基于 Node.js 的前端自动化任务运行器，主要用于自动化执行项目开发中重复性高的任务，例如文件压缩、代码编译、语法检查、单元测试、文件合并等。通过配置简洁的任务…...

编程新知 2025/7/7 18:23:43

打手机检测算法AI智能分析网关V4守护公共/工业/医疗等多场景安全应用

一、方案背景在现代生产与生活场景中，如工厂高危作业区、医院手术室、公共场景等，人员违规打手机的行为潜藏着巨大风险。传统依靠人工巡查的监管方式，存在效率低、覆盖面不足、判断主观性强等问题，难以满足对人员打手机行为精…...

编程新知 2025/7/7 10:57:34

MySQL 主从同步异常处理

阅读原文：https://www.xiaozaoshu.top/articles/mysql-m-s-update-pk MySQL 做双主，遇到的这个错误： Could not execute Update_rows event on table ... Error_code: 1032是 MySQL 主从复制时的经典错误之一，通常表示&#xff…...

编程新知 2025/7/7 4:07:39

Ref

相关文章：