- 数据增强
- 计算机视觉
- 图像处理
- 机器学习
【免费下载链接】imgaug
Image augmentation for machine learning experiments.
imgaug 是一个面向机器学习(尤其是卷积神经网络)实验的 Python 图像增强库,其核心思路是把一组输入图像转换成更大、更多样的一组"微调"图像,从而扩充训练数据、缓解过拟合。本文以仓库根目录的 README.md 为骨架,结合 imgaug/augmenters/meta.py 等源码实现,系统讲解 imgaug 的特性与安装方式、13 大模块的增强器全览、以及从"简单训练流程"到"复杂增强流水线"、从图像到关键点/边界框/热力图等各类标注数据对齐增强的完整实战方案。
核心特性:一张图看懂 imgaug 能做什么
按 README.md 的 Features 章节,imgaug 的核心能力可以归纳为六个方面:
- 丰富的增强技术:包括仿射变换(affine)、透视变换(perspective)、对比度变化(contrast)、高斯噪声(gaussian noise)、区域丢弃(dropout)、色相/饱和度调整(hue/saturation)、裁剪/填充(crop/pad)、模糊(blur)等,覆盖几何、颜色、噪声、风格化等多个维度;
- 面向高性能优化:增强流程针对批量数据处理做了性能优化,0.4.0 起后端改为逐批次(batchwise)增强;
- 多类型标注数据支持:
- 图像(uint8 完整支持,其他 dtype 支持情况因增强器而异);
- 热力图 Heatmaps(float32)、分割图 Segmentation Maps(int)、掩码 Masks(bool)——它们可以与图像尺寸不一致,例如裁剪时无需额外代码即可自动按比例处理;
- 关键点/地标 Keypoints(int/float 坐标);
- 边界框 Bounding Boxes(int/float 坐标);
- 多边形 Polygons(int/float 坐标);
- 线串 Line Strings(int/float 坐标);
- 随机值的自动对齐:例如从
uniform(-10°, 45°)采样一个旋转角,图像和叠加其上的分割图会自动使用同一个采样值,零额外代码; - 概率分布作为参数:例如从
uniform(-10°, 45°)采样旋转角度,甚至支持ABS(N(0, 20.0))*(1+B(1.0, 1.0))这类由绝对值函数ABS(.)、高斯分布N(.)、Beta 分布B(.)组合而成的复杂表达式; - 丰富的辅助函数:绘制热力图、分割图、关键点、边界框等;缩放分割图、对图像/地图做平均池化或最大池化、将图像 pad 到指定宽高比(如正方形);把关键点转换为距离图、从图像中提取边界框内的像素、把多边形裁剪到图像平面等;
- 多 CPU 核心并行增强:支持在后台进程中对多个 batch 进行并行增强。
安装与依赖环境
Anaconda 安装
README 提供官方 Anaconda 安装方式,通过 conda-forge 渠道安装:
conda config --add channels conda-forge conda install imgaug卸载则执行conda remove imgaug。
pip 安装
通过 PyPI 安装(可能滞后于 GitHub 源码版本):
pip install imgaug或直接从 GitHub 安装最新版:
pip install git+https://github.com/aleju/imgaug.git卸载执行pip uninstall imgaug。
版本与运行时要求
当前仓库为0.4.0版本(见 setup.py),README 声明支持 Python 2.7 与 3.4+;setup.py 的 classifiers 进一步确认支持 Python 2.7 及 3.4~3.8。其核心运行时依赖(setup.py)为:
| 依赖包 | 版本约束 | 用途 |
|---|---|---|
| six | 无 | Python 2/3 兼容层 |
| numpy | >=1.15 | 数值数组基础 |
| scipy | 无 | 滤波/插值等科学计算 |
| Pillow | 无 | 图像读写与 PIL 风格增强 |
| matplotlib | 无 | 绘图与可视化 |
| scikit-image | >=0.14.2 | 仿射/透视等几何变换底层 |
| opencv-python-headless | 无 | OpenCV 运算(可替换) |
| imageio | 无 | 图像/视频读写 |
| Shapely | 无 | 多边形几何运算 |
注意一个细节:opencv-python-headless存在三个可替换项(opencv-python、opencv-contrib-python、opencv-contrib-python-headless)。setup.py 的check_alternative_installation()会在安装时检测用户环境里是否已存在这些替代包,若已安装则不再重复装 OpenCV,避免同一库的多重安装冲突。
文档资源
README 推荐了两类学习资源:官方 Jupyter Notebook(涵盖图像加载与增强、多核增强,以及关键点、边界框、多边形、线串、热力图、分割图的增强操作)和 ReadTheDocs 文档(快速上手示例、全部增强器总览、API 参考)。仓库内另有 checks/ 目录,包含数十个可直接运行的示例脚本(如 check_affine.py、check_some_of.py、check_multicore_pool.py),是本地验证增强效果、学习参数用法的快捷入口。
版本演进脉络
README 的 Recent Changes 章节与仓库 changelogs/ 目录记录了各版本的关键变化:
- 0.4.0:新增多个增强器;增强后端改为 batchwise 逐批次增强;支持 numpy 1.18 与 Python 3.8。这一"batchwise 后端"在源码中有直接印证——imgaug/augmenters/meta.py 的
_augment_batch_方法注释明确标注 "Added in 0.4.0",它统一处理同一 batch 内图像、关键点等多列数据的对齐采样; - 0.3.0:重构分割图增强;适配 numpy 1.17+ 的随机数采样 API;新增多个增强器;
- 0.2.9:新增多边形增强、线串增强,简化增强接口;
- 0.2.8:改进性能、dtype 支持与多核增强。
增强器全景总览(按模块分类)
README 的 Example Images 章节按模块列出了绝大多数增强器。其中形如(a, b)的参数值表示从区间[a, b]中随机均匀采样。线串(Line Strings)几乎被所有增强器支持,只是在该章节中未单独可视化。下表汇总各模块及其代表增强器,对应源码模块位于 imgaug/augmenters/ 下:
| 模块(源码文件) | 代表增强器 | 其他可用增强器 |
|---|---|---|
| meta(meta.py) | Identity、ChannelShuffle | Sequential、SomeOf、OneOf、Sometimes、WithChannels、Lambda、AssertLambda、AssertShape、RemoveCBAsByOutOfImageFraction、ClipCBAsToImagePlanes |
| arithmetic(arithmetic.py) | Add、AdditiveGaussianNoise、Multiply、Cutout、Dropout、CoarseDropout、Dropout2d、SaltAndPepper、CoarseSaltAndPepper、Invert、Solarize、JpegCompression | AddElementwise、AdditiveLaplaceNoise、AdditivePoissonNoise、MultiplyElementwise、TotalDropout、ReplaceElementwise、ImpulseNoise、Salt、Pepper、CoarseSalt、CoarsePepper |
| artistic(artistic.py) | Cartoon | — |
| blend(blend.py) | BlendAlpha、BlendAlphaSimplexNoise、BlendAlphaFrequencyNoise、BlendAlphaSomeColors、BlendAlphaRegularGrid | BlendAlphaMask、BlendAlphaElementwise、BlendAlphaVerticalLinearGradient、BlendAlphaHorizontalLinearGradient、BlendAlphaSegMapClassIds、BlendAlphaBoundingBoxes、BlendAlphaCheckerboard,及 SomeColorsMaskGen、RegularGridMaskGen、CheckerboardMaskGen、InvertMaskGen 等 MaskGen |
| blur(blur.py) | GaussianBlur、AverageBlur、MedianBlur、BilateralBlur、MotionBlur、MeanShiftBlur | — |
| collections(collections.py) | RandAugment | — |
| color(color.py) | MultiplyAndAddToBrightness、MultiplyHueAndSaturation、MultiplyHue、MultiplySaturation、AddToHueAndSaturation、Grayscale、RemoveSaturation、ChangeColorTemperature、KMeansColorQuantization、UniformColorQuantization | WithColorspace、WithBrightnessChannels、MultiplyBrightness、AddToBrightness、WithHueAndSaturation、AddToHue、AddToSaturation、ChangeColorspace、Posterize |
| contrast(contrast.py) | GammaContrast、SigmoidContrast、LogContrast、LinearContrast、HistogramEqualization、AllChannelsHistogramEqualization、AllChannelsCLAHE、CLAHE | Equalize |
| convolutional(convolutional.py) | Sharpen、Emboss、EdgeDetect、DirectedEdgeDetect | Convolve |
| debug(debug.py) | — | SaveDebugImageEveryNBatches |
| edges(edges.py) | Canny | — |
| flip(flip.py) | Fliplr、Flipud | HorizontalFlip、VerticalFlip |
| geometric(geometric.py) | Affine(含 Modes/cval 变体)、PiecewiseAffine、PerspectiveTransform、ElasticTransformation、Rot90、WithPolarWarping、Jigsaw | ScaleX、ScaleY、TranslateX、TranslateY、Rotate |
| imgcorruptlike(imgcorruptlike.py) | GlassBlur、DefocusBlur、ZoomBlur、Snow、Spatter | GaussianNoise、ShotNoise、ImpulseNoise、SpeckleNoise、Fog、Frost、Contrast、Brightness、Saturate、JpegCompression、Pixelate、ElasticTransform |
| pillike(pillike.py) | Autocontrast、EnhanceColor、EnhanceSharpness、FilterEdgeEnhanceMore、FilterContour | Solarize、Posterize、Equalize、EnhanceContrast、EnhanceBrightness、FilterBlur、FilterSmooth、FilterSmoothMore、FilterEdgeEnhance、FilterFindEdges、FilterEmboss、FilterSharpen、FilterDetail、Affine |
| pooling(pooling.py) | AveragePooling、MaxPooling、MinPooling、MedianPooling | — |
| segmentation(segmentation.py) | Superpixels、UniformVoronoi、RegularGridVoronoi | Voronoi、RelativeRegularGridVoronoi,及 RegularGridPointsSampler、UniformPointsSampler、DropoutPointsSampler、SubsamplingPointsSampler 等 PointsSampler |
| size(size.py) | CropAndPad、Crop、Pad、PadToFixedSize、CropToFixedSize | Resize、CropToMultiplesOf、PadToMultiplesOf、CropToPowersOf、PadToPowersOf、CropToAspectRatio、PadToAspectRatio、CropToSquare、PadToSquare,以及对应的 Center 系列与 KeepSizeByResize |
| weather(weather.py) | FastSnowyLandscape、Clouds、Fog、Snowflakes、Rain | CloudLayer、SnowflakesLayer、RainLayer |
实战一:标准训练流程中的简单增强
README 的第一个代码示例演示了最常见的机器学习训练场景——每个 batch 依次做随机裁剪、水平翻转、高斯模糊。其要点是:图像输入约定为(N, height, width, channels)的 numpy 数组,或不同尺寸(height, width, channels)数组组成的列表;做颜色空间类增强时应使用 RGB(cv2.imread()返回的是 BGR);图像通常使用取值 0~255 的uint8。
import numpy as np import imgaug.augmenters as iaa def load_batch(batch_idx): # dummy function, implement this # Return a numpy array of shape (N, height, width, #channels) # or a list of (height, width, #channels) arrays (may have different image # sizes). # Images should be in RGB for colorspace augmentations. # (cv2.imread() returns BGR!) # Images should usually be in uint8 with values from 0-255. return np.zeros((128, 32, 32, 3), dtype=np.uint8) + (batch_idx % 255) def train_on_images(images): # dummy function, implement this pass # Pipeline: # (1) Crop images from each side by 1-16px, do not resize the results # images back to the input size. Keep them at the cropped size. # (2) Horizontally flip 50% of the images. # (3) Blur images using a gaussian kernel with sigma between 0.0 and 3.0. seq = iaa.Sequential([ iaa.Crop(px=(1, 16), keep_size=False), iaa.Fliplr(0.5), iaa.GaussianBlur(sigma=(0, 3.0)) ]) for batch_idx in range(100): images = load_batch(batch_idx) images_aug = seq(images=images) # done by the library train_on_images(images_aug)其中Crop(px=(1, 16), keep_size=False)表示从每侧随机裁剪 1~16 像素且不把结果缩放回原尺寸;Fliplr(0.5)表示 50% 概率水平翻转;GaussianBlur(sigma=(0, 3.0))使用 0.0~3.0 的高斯核 sigma 值模糊。
从源码层面看,seq(images=images)等价于调用 imgaug/augmenters/meta.py 的augment()方法,它会将输入包装成UnnormalizedBatch,再经augment_batch进入 meta.py 的_augment_batch_统一处理;而Sequential本身就是list的子类(meta.py),它会依次把每个子增强器应用到数据上,即"第二个增强器接收的是已经过第一个增强器处理的输入"。
实战二:非常复杂的增强流水线
README 给出了用于生成首页效果图的"重度"增强流水线,是组合各类增强器的最佳示范。先定义sometimes辅助函数,让指定增强器只在 50% 的样本上生效:
import numpy as np import imgaug as ia import imgaug.augmenters as iaa # random example images images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) # Sometimes(0.5, ...) applies the given augmenter in 50% of all cases, # e.g. Sometimes(0.5, GaussianBlur(0.3)) would blur roughly every second image. sometimes = lambda aug: iaa.Sometimes(0.5, aug) seq = iaa.Sequential( [ # apply the following augmenters to most images iaa.Fliplr(0.5), # horizontally flip 50% of all images iaa.Flipud(0.2), # vertically flip 20% of all images # crop images by -5% to 10% of their height/width sometimes(iaa.CropAndPad( percent=(-0.05, 0.1), pad_mode=ia.ALL, pad_cval=(0, 255) )), sometimes(iaa.Affine( scale={"x": (0.8, 1.2), "y": (0.8, 1.2)}, # scale images to 80-120% of their size, individually per axis translate_percent={"x": (-0.2, 0.2), "y": (-0.2, 0.2)}, # translate by -20 to +20 percent (per axis) rotate=(-45, 45), # rotate by -45 to +45 degrees shear=(-16, 16), # shear by -16 to +16 degrees order=[0, 1], # use nearest neighbour or bilinear interpolation (fast) cval=(0, 255), # if mode is constant, use a cval between 0 and 255 mode=ia.ALL # use any of scikit-image's warping modes )), # execute 0 to 5 of the following (less important) augmenters per image # don't execute all of them, as that would often be way too strong iaa.SomeOf((0, 5), [ sometimes(iaa.Superpixels(p_replace=(0, 1.0), n_segments=(20, 200))), # convert images into their superpixel representation iaa.OneOf([ iaa.GaussianBlur((0, 3.0)), # blur images with a sigma between 0 and 3.0 iaa.AverageBlur(k=(2, 7)), # blur image using local means with kernel sizes between 2 and 7 iaa.MedianBlur(k=(3, 11)), # blur image using local medians with kernel sizes between 2 and 7 ]), iaa.Sharpen(alpha=(0, 1.0), lightness=(0.75, 1.5)), # sharpen images iaa.Emboss(alpha=(0, 1.0), strength=(0, 2.0)), # emboss images # search either for all edges or for directed edges, # blend the result with the original image using a blobby mask iaa.SimplexNoiseAlpha(iaa.OneOf([ iaa.EdgeDetect(alpha=(0.5, 1.0)), iaa.DirectedEdgeDetect(alpha=(0.5, 1.0), direction=(0.0, 1.0)), ])), iaa.AdditiveGaussianNoise(loc=0, scale=(0.0, 0.05*255), per_channel=0.5), # add gaussian noise to images iaa.OneOf([ iaa.Dropout((0.01, 0.1), per_channel=0.5), # randomly remove up to 10% of the pixels iaa.CoarseDropout((0.03, 0.15), size_percent=(0.02, 0.05), per_channel=0.2), ]), iaa.Invert(0.05, per_channel=True), # invert color channels iaa.Add((-10, 10), per_channel=0.5), # change brightness of images (by -10 to 10 of original value) iaa.AddToHueAndSaturation((-20, 20)), # change hue and saturation # either change the brightness of the whole image (sometimes # per channel) or change the brightness of subareas iaa.OneOf([ iaa.Multiply((0.5, 1.5), per_channel=0.5), iaa.FrequencyNoiseAlpha( exponent=(-4, 0), first=iaa.Multiply((0.5, 1.5), per_channel=True), second=iaa.LinearContrast((0.5, 2.0)) ) ]), iaa.LinearContrast((0.5, 2.0), per_channel=0.5), # improve or worsen the contrast iaa.Grayscale(alpha=(0.0, 1.0)), sometimes(iaa.ElasticTransformation(alpha=(0.5, 3.5), sigma=0.25)), # move pixels locally around (with random strengths) sometimes(iaa.PiecewiseAffine(scale=(0.01, 0.05))), # sometimes move parts of the image around sometimes(iaa.PerspectiveTransform(scale=(0.01, 0.1))) ], random_order=True ) ], random_order=True ) images_aug = seq(images=images)这段流水线集中体现了 README Features 章节所述的几大设计:
Sometimes(0.5, aug)让增强器以 50% 概率生效;SomeOf((0, 5), [...], random_order=True)每张图随机执行列表中的 0~5 个子增强器且顺序随机——避免一次施加全部增强导致过度失真;OneOf([...])每次只从列表中随机选一个执行,例如三种模糊效果二选一/三选一;per_channel=0.5表示 50% 情况下按"整张图采样一个值",其余情况按"每个通道各自采样一个值";ia.ALL、iaa.Affine(mode=ia.ALL)等表示从全部合法取值中随机选择(例如 scikit-image 的全部 warp 模式)。
这些容器增强器在 imgaug/augmenters/meta.py 中均有对应实现:Sequential(L3006)顺序应用子增强器并支持random_order;SomeOf(L3188)随机挑选 n 个子增强器;OneOf(L3470)是SomeOf的特例,每次恰好激活一个子增强器;Sometimes(L3539)支持then_list/else_list两个分支。Sequential的 docstring 还特别说明random_order=True时子增强器的顺序会在每个 batch 随机采样一次,能显著扩大增强空间。
实战三:图像与关键点/地标的对齐增强
目标检测、姿态估计等任务通常需要同时增强图像与其上的关键点。README 示例中,两张测试图在(64, 64)处标记白色像素,关键点分别为第一张 1 个点、第二张 3 个点,增强序列为高斯噪声 + 沿 x 轴平移 1~5 像素:
import numpy as np import imgaug.augmenters as iaa images = np.zeros((2, 128, 128, 3), dtype=np.uint8) # two example images images[:, 64, 64, :] = 255 points = [ [(10.5, 20.5)], # points on first image [(50.5, 50.5), (60.5, 60.5), (70.5, 70.5)] # points on second image ] seq = iaa.Sequential([ iaa.AdditiveGaussianNoise(scale=0.05*255), iaa.Affine(translate_px={"x": (1, 5)}) ]) # augment keypoints and images images_aug, points_aug = seq(images=images, keypoints=points) print("Image 1 center", np.argmax(images_aug[0, 64, 64:64+6, 0])) print("Image 2 center", np.argmax(images_aug[1, 64, 64:64+6, 0])) print("Points 1", points_aug[0]) print("Points 2", points_aug[1])README 特别强调:imgaug 中所有坐标都是亚像素精度(subpixel-accurate),因此x=0.5, y=0.5表示左上角像素的中心。这条约定同样适用于下文的所有坐标类数据。
实战四:边界框、多边形与线串
三种坐标类标注数据的增强写法高度一致,均由 "图像 + 标注列表" 构成输入,seq(images=..., xxx=...)返回增强后的两者。
边界框(坐标形式为x1, y1, x2, y2):
import numpy as np import imgaug as ia import imgaug.augmenters as iaa images = np.zeros((2, 128, 128, 3), dtype=np.uint8) # two example images images[:, 64, 64, :] = 255 bbs = [ [ia.BoundingBox(x1=10.5, y1=15.5, x2=30.5, y2=50.5)], [ia.BoundingBox(x1=10.5, y1=20.5, x2=50.5, y2=50.5), ia.BoundingBox(x1=40.5, y1=75.5, x2=70.5, y2=100.5)] ] seq = iaa.Sequential([ iaa.AdditiveGaussianNoise(scale=0.05*255), iaa.Affine(translate_px={"x": (1, 5)}) ]) images_aug, bbs_aug = seq(images=images, bounding_boxes=bbs)多边形(每个多边形由 3 个以上顶点定义):
import numpy as np import imgaug as ia import imgaug.augmenters as iaa images = np.zeros((2, 128, 128, 3), dtype=np.uint8) # two example images images[:, 64, 64, :] = 255 polygons = [ [ia.Polygon([(10.5, 10.5), (50.5, 10.5), (50.5, 50.5)])], [ia.Polygon([(0.0, 64.5), (64.5, 0.0), (128.0, 128.0), (64.5, 128.0)])] ] seq = iaa.Sequential([ iaa.AdditiveGaussianNoise(scale=0.05*255), iaa.Affine(translate_px={"x": (1, 5)}) ]) images_aug, polygons_aug = seq(images=images, polygons=polygons)线串(与多边形类似,但不闭合、可自交、无内部面积,适合车道线、骨骼等场景):
import numpy as np import imgaug as ia import imgaug.augmenters as iaa images = np.zeros((2, 128, 128, 3), dtype=np.uint8) # two example images images[:, 64, 64, :] = 255 ls = [ [ia.LineString([(10.5, 10.5), (50.5, 10.5), (50.5, 50.5)])], [ia.LineString([(0.0, 64.5), (64.5, 0.0), (128.0, 128.0), (64.5, 128.0), (128.0, 0.0)])] ] seq = iaa.Sequential([ iaa.AdditiveGaussianNoise(scale=0.05*255), iaa.Affine(translate_px={"x": (1, 5)}) ]) images_aug, ls_aug = seq(images=images, line_strings=ls)实战五:热力图与分割图(不同尺寸自动对齐)
热力图是取值 0.0~1.0 的稠密 float 数组,常用于训练人脸关键点定位等模型;分割图则为int32稠密数组。两者都可以与图像尺寸不同——README 示例中热力图/分割图是64x64,而图像是128x128,imgaug 会自动处理这种差异,例如图像每侧裁剪 10 像素时,热力图只裁剪一半。
热力图:
import numpy as np import imgaug.augmenters as iaa # Standard scenario: You have N RGB-images and additionally 21 heatmaps per # image. You want to augment each image and its heatmaps identically. images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) heatmaps = np.random.random(size=(16, 64, 64, 1)).astype(np.float32) seq = iaa.Sequential([ iaa.GaussianBlur((0, 3.0)), iaa.Affine(translate_px={"x": (-40, 40)}), iaa.Crop(px=(0, 10)) ]) images_aug, heatmaps_aug = seq(images=images, heatmaps=heatmaps)分割图(缩放等操作会自动使用最近邻插值,避免产生非整数类别):
import numpy as np import imgaug.augmenters as iaa # Standard scenario: You have N=16 RGB-images and additionally one segmentation # map per image. You want to augment each image and its heatmaps identically. images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) segmaps = np.random.randint(0, 10, size=(16, 64, 64, 1), dtype=np.int32) seq = iaa.Sequential([ iaa.GaussianBlur((0, 3.0)), iaa.Affine(translate_px={"x": (-40, 40)}), iaa.Crop(px=(0, 10)) ]) images_aug, segmaps_aug = seq(images=images, segmentation_maps=segmaps)多列数据(如图像 + 边界框)在同 batch 内使用相同采样值的对齐机制,在 meta.py 的_augment_batch_中有明确实现:当 batch 含多列数据时,会自动进入该 batch 内的确定性采样模式,避免各数据类型拿到不同的随机样本。
实战六:结果可视化
可视化增强后的图像,使用show_grid一次性铺开rows × cols个增强结果(README 示例生成 8×8 网格,对两张输入图施加相同增强):
import numpy as np import imgaug.augmenters as iaa images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) seq = iaa.Sequential([iaa.Fliplr(0.5), iaa.GaussianBlur((0, 3.0))]) # Show an image with 8*8 augmented versions of image 0 and 8*8 augmented # versions of image 1. Identical augmentations will be applied to # image 0 and 1. seq.show_grid([images[0], images[1]], cols=8, rows=8)可视化非图像数据(README 的辅助函数示例,涵盖关键点、边界框、多边形、热力图的draw_on_image):
import numpy as np import imgaug as ia image = np.zeros((64, 64, 3), dtype=np.uint8) # points kps = [ia.Keypoint(x=10.5, y=20.5), ia.Keypoint(x=60.5, y=60.5)] kpsoi = ia.KeypointsOnImage(kps, shape=image.shape) image_with_kps = kpsoi.draw_on_image(image, size=7, color=(0, 0, 255)) ia.imshow(image_with_kps) # bbs bbsoi = ia.BoundingBoxesOnImage([ ia.BoundingBox(x1=10.5, y1=20.5, x2=50.5, y2=30.5) ], shape=image.shape) image_with_bbs = bbsoi.draw_on_image(image) image_with_bbs = ia.BoundingBox( x1=50.5, y1=10.5, x2=100.5, y2=16.5 ).draw_on_image(image_with_bbs, color=(255, 0, 0), size=3) ia.imshow(image_with_bbs) # polygons psoi = ia.PolygonsOnImage([ ia.Polygon([(10.5, 20.5), (50.5, 30.5), (10.5, 50.5)]) ], shape=image.shape) image_with_polys = psoi.draw_on_image( image, alpha_points=0, alpha_face=0.5, color_lines=(255, 0, 0)) ia.imshow(image_with_polys) # heatmaps hms = ia.HeatmapsOnImage(np.random.random(size=(32, 32, 1)).astype(np.float32), shape=image.shape) image_with_hms = hms.draw_on_image(image) ia.imshow(image_with_hms)LineStrings 与分割图支持与此类似的方法。这些辅助函数对调试数据增强前后的标注对齐非常有用。
实战七:一次性使用增强器
虽然接口设计鼓励复用增强器实例,但也可以即用即弃,实例化开销通常可忽略:
from imgaug import augmenters as iaa import numpy as np images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) # always horizontally flip each input image images_aug = iaa.Fliplr(1.0)(images=images) # vertically flip each input image with 90% probability images_aug = iaa.Flipud(0.9)(images=images) # blur 50% of all images using a gaussian kernel with a sigma of 3.0 images_aug = iaa.Sometimes(0.5, iaa.GaussianBlur(3.0))(images=images)实战八:多核后台批量增强
当数据量很大时,可以把增强放到后台进程执行。核心 API 是augment_batches(batches, background=True),其中batches为 imgaug.augmentables.batches.UnnormalizedBatch 或Batch的列表/生成器。README 示例用同一张图构造 10 个 batch、每 batch 32 张图,并用draw_grid展示结果:
import skimage.data import imgaug as ia import imgaug.augmenters as iaa from imgaug.augmentables.batches import UnnormalizedBatch # Number of batches and batch size for this example nb_batches = 10 batch_size = 32 # Example augmentation sequence to run in the background augseq = iaa.Sequential([ iaa.Fliplr(0.5), iaa.CoarseDropout(p=0.1, size_percent=0.1) ]) # For simplicity, we use the same image here many times astronaut = skimage.data.astronaut() astronaut = ia.imresize_single_image(astronaut, (64, 64)) # Make batches out of the example image (here: 10 batches, each 32 times # the example image) batches = [] for _ in range(nb_batches): batches.append(UnnormalizedBatch(images=[astronaut] * batch_size)) # Show the augmented images. # Note that augment_batches() returns a generator. for images_aug in augseq.augment_batches(batches, background=True): ia.imshow(ia.draw_grid(images_aug.images_aug, cols=8))从 meta.py 的augment_batches实现可以看到几个关键行为:
- 该方法**产出(yield)**增强后的 batch,而不是一次性返回完整列表,更适合训练循环的流式消费;
background=True时会基于imgaug.multicore.Pool启动后台进程池,默认使用所有可用逻辑 CPU 核,输出缓冲区大小为C*10(C为逻辑核数);- 多核模式按"batch 粒度"分发数据,不会把单个 batch 内的数据拆分到不同核,因此对单 batch 使用
background=True没有意义;且后台模式下hooks不可用(涉及函数序列化); - 若需要更精细的控制(设置种子、指定 CPU 核数、限制内存),README 指向
Augmenter.pool()与imgaug.multicore.Pool,对应实现见 imgaug/multicore.py。
实战九:概率分布作为参数
多数增强器的参数支持两种快捷写法:元组(a, b)表示uniform(a, b)均匀分布,列表[a, b, c]表示从给定集合中随机挑一个。需要更复杂分布(高斯、截断高斯、泊松等)时,可用 imgaug/parameters.py 中的随机参数:
import numpy as np from imgaug import augmenters as iaa from imgaug import parameters as iap images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) # Blur by a value sigma which is sampled from a uniform distribution # of range 10.1 <= x < 13.0. # The convenience shortcut for this is: GaussianBlur((10.1, 13.0)) blurer = iaa.GaussianBlur(10 + iap.Uniform(0.1, 3.0)) images_aug = blurer(images=images) # Blur by a value sigma which is sampled from a gaussian distribution # N(1.0, 0.1), i.e. sample a value that is usually around 1.0. # Clip the resulting value so that it never gets below 0.1 or above 3.0. blurer = iaa.GaussianBlur(iap.Clip(iap.Normal(1.0, 0.1), 0.1, 3.0)) images_aug = blurer(images=images)库中还提供了截断高斯分布、泊松分布、Beta 分布等更多概率分布(README Features 章节中的ABS(N(0, 20.0))*(1+B(1.0, 1.0))即此类组合的典型示例)。
实战十:按通道增强(WithChannels)
某些场景下只想增强图像的指定通道(如 R、G 通道)。WithChannels正是为此设计:
import numpy as np import imgaug.augmenters as iaa # fake RGB images images = np.random.randint(0, 255, (16, 128, 128, 3), dtype=np.uint8) # add a random value from the range (-30, 30) to the first two channels of # input images (e.g. to the R and G channels) aug = iaa.WithChannels( channels=[0, 1], children=iaa.Add((-30, 30)) ) images_aug = aug(images=images)源码层面,meta.py 的WithChannels实现会先把图像缩减到指定通道,在子增强器上完成处理后,再把未增强的其他通道替换回原值——这正是 README Features 中"Easy to apply augmentations only to some images/channels"的具体落地。
如何在研究中使用与引用
若该库对研究有帮助,README 提供了 BibTeX 引用条目(作者为 Alexander Jung 等贡献者,年 2020)。此外,仓库根目录 CHANGELOG.md 与 changelogs/ 目录记录了自 0.2.8 以来的全部变更细节(新增、修改、废弃、修复、重构),升级版本前建议查阅对应版本的变更文档以确认接口变化(例如 0.4.0 中random_state/deterministic参数已标记为废弃,推荐改用seed与to_deterministic())。
综上,从简单的三行增强序列,到覆盖十几个模块的重度流水线,再到多核后台批量增强,imgaug 0.4.0 提供了一整套开箱即用的图像与标注数据对齐增强方案。读者可以以本文的示例为起点,结合 test/ 下的测试用例与 checks/ 下的可运行脚本,进一步验证每个增强器在不同参数下的实际表现。
- 数据增强
- 计算机视觉
- 图像处理
- 机器学习
【免费下载链接】imgaug
Image augmentation for machine learning experiments.
相关推荐
图像增强库ImgAug安装与使用指南
图像增强库ImgAug安装与使用指南 一、项目介绍 ImgAug , 即 Image Augmentation , 是一个专用于图像增强的Python库,在深度
数据增强计算机视觉图像处理机器学习图像增强神器 imgaug:机器学习数据增强的终极指南 🚀
图像增强神器 imgaug:机器学习数据增强的终极指南 🚀 在机器学习特别是计算机视觉项目中,数据增强是提升模型泛化能力的关键技术。imgaug 是一个功能强
数据增强计算机视觉图像处理机器学习终极图像增强指南:用imgaug实现虚实融合的增强现实技术 🚀
终极图像增强指南:用imgaug实现虚实融合的增强现实技术 🚀 在机器学习的世界里,数据是燃料,而 图像增强技术 就是让燃料更高效、更丰富的秘密武器!今天我要
数据增强计算机视觉图像处理机器学习
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考