数据填充

数据填充模块是一个面向 Odoo 数据库的合成数据生成框架。它遵循声明式 Blueprint(蓝图) 模式:你在 XML 或 JSON 中描述 要创建的数据,系统便会大规模生成记录,并支持并行执行、统计分布以及生成目标之间的依赖关系。

典型使用场景:

  • 性能测试 —— 生成数千条记录,以压力测试查询、视图和报表。

  • 演示环境 —— 随模块附带一份看起来真实的数据集。

  • 开发 —— 快速填充本地数据库,以便开发依赖既有数据的功能。

另请参阅

duplicate - 复制记录 提供了一个更简单的工具,用于批量复制 已有 记录。

安装

  1. 在你的数据库上安装 populate Odoo 模块。

  2. (可选)安装 Faker 库,以启用 fake.* 生成器:

    $ pip install -r odoo/addons/populate/requirements.txt
    

重要

安装了新提供蓝图的模块后,你必须 升级 populate 模块 ,以便发现其蓝图并将其加载到数据库中:

$ odoo-bin -d <database> -u populate

CLI 命令

$ odoo-bin populate -d <database> -b <blueprint>
-d <database>, --database <database>

目标数据库(必填)。

-b <blueprint>, --blueprint <blueprint>

蓝图名称或完整 xmlid(必填,使用 --resume 时除外)。

--seed <seed>

随机数生成器的种子。若省略,将随机选取一个种子。提供相同的种子可保证结果可复现(确定性生成)。

--scale <factor>

按此因子缩放蓝图中所有记录数量。默认值:1。

-j <workers>, --jobs <workers>

并行工作进程数。使用 auto 可启用所有可用的 CPU 线程。默认值:1。

--resume [session_id]

恢复一个中断的会话。不带参数时,恢复最近一次未完成的会话;提供会话 ID 时,恢复该特定会话。

--profile

为本次运行中的每个可执行 populate 作业保存 profiler 跟踪。

Example

# Run a blueprint at 10x scale using all CPU cores
$ odoo-bin populate -d mydb -b project.fake_project_demo --scale 10 -j auto

# Run with a fixed seed for reproducibility
$ odoo-bin populate -d mydb -b my_module.my_blueprint --seed 42

# Resume the last interrupted session
$ odoo-bin populate -d mydb --resume

# Resume a specific session by ID
$ odoo-bin populate -d mydb --resume 7

# Save profiler traces for the generated workload
$ odoo-bin populate -d mydb -b project.fake_project_demo --profile

对 populate 运行进行性能剖析

使用 --profile 测量一次 populate 运行的运行时开销。该命令会创建 ir.profile 条目,并按以 populate 会话和蓝图命名的一个 profiler 会话进行分组。

该选项仅对当前命令调用生效,因此同一个中断的会话可稍后带 profiling 或不带 profiling 恢复:

$ odoo-bin populate -d mydb --resume 7 --profile

profiling 在单工作线程和多工作线程模式下均可用。每个可执行作业都会创建自己的 profile 条目,包括大作业被拆分并行执行时产生的子作业。仅协调拆分后子作业的 Planner 作业本身不创建或更新记录,因此不会被 profiling。失败的可执行作业尝试同样会创建 profile 条目,因此它们的跟踪信息仍可用于调试分析。

蓝图

蓝图(Blueprint) 是一条 populate.blueprint 记录,声明式地描述了要创建的数据。蓝图通常随模块的 populate/ 数据文件夹一起发布,并在升级 populate 模块时被自动加载。

如果某模块的 populate/ 文件夹是有效的 Python 包(包含 __init__.py),则其代码会被导入,从而允许该模块注册自定义生成器。

蓝图可以以 XML、JSON 或两者兼有来定义。若同一条记录上同时设置了 definition_xml 和 definition_json,则以 XML 定义为准。

操作块

一个蓝图定义是一个有序的块(block)列表。用 <create/> 创建记录,用 <write/> 更新目标记录,用 <function/> 在目标记录上调用方法。字段通过 ORM 的 create 或 write 调用持久化,而值是局部变量,可在同一块中稍后的字段、值和函数参数里复用。

Example

<create model="res.partner" count="500" id="my_partners">
    <value name="email_domain" eval="'example.com'"/>
    <field name="name" generator="fake.company"/>
    <field name="email"
           eval="name.lower().replace(' ', '-') + '@' + email_domain"/>
    <field name="active" eval="True"/>
</create>

<write model="res.partner" ref="my_partners" batched="True">
    <value name="suffix" eval="'imported by populate'"/>
    <field name="comment" eval="suffix"/>
</write>

<function model="res.partner" name="message_subscribe"
          ref="my_partners" batched="True">
    <value name="current_partner" eval="env.user.partner_id.id"/>
    <arg eval="[current_partner]"/>
</function>
model (必填)

Odoo 模型的技术名称,例如 res.partner。

name (function 必填)

要调用的方法名称。

count (create 必填)

要创建的记录数量。

id

create 块的引用标签。后续块可通过 ref 定位这些记录。

ref

对于 write 和 function 块:引用之前创建的一个批次(其 id)。

domain

对于 write 和 function 块:用于选择目标记录的 ORM 域。若既未提供 ref 也未提供 domain,则以该模型的所有记录为目标。create 块无法定义顶层 domain,因为它们创建新记录而非定位已有记录。

scale

True (默认)或 False。是否对当前块的 count 应用 --scale 因子。

parallel

True (默认)或 False。是否可将该作业拆分到并行工作线程中执行。当模型的约束要求顺序写入时,请设为 False。

batched

仅适用于 write 和 function 块。设为 True 时,每个可执行作业生成一个值集并执行一次写入或函数调用。默认为 False,此时会为每条目标记录分别生成值并执行操作。参见 写入任务 和 函数任务。

context

一个 Python 字典字面量,将被合并到 create、write 或函数调用的 ORM 上下文中。

重要

操作块按文档顺序执行。通过 ref 引用其他块的块,只能指向在蓝图中 更早 定义的块。请先定义主数据,再定义依赖于它的记录:

Example

<!-- 1. Stage definitions (master data) -->
<create model="project.task.type" count="8" id="task_types" scale="False">
    <field name="name" generator="fake.bs"/>
</create>

<!-- 2. Projects (reference stages) -->
<create model="project.project" count="120" id="projects">
    <field name="type_ids" ref="task_types" count="8"/>
</create>

<!-- 3. Tasks (reference projects and stages) -->
<create model="project.task" count="10000" id="tasks">
    <field name="project_id" ref="projects"/>
    <field name="stage_id"   ref="task_types"/>
</create>

JSON 格式

JSON 格式与 XML 结构相对应。必填的 operation 键对应 XML 中的操作元素名称(create、write 或 function)。顶层数组对应有序的块列表:

Example

[
    {
        "operation": "create",
        "model": "res.partner",
        "count": 500,
        "id": "my_partners",
        "values": {
            "email_domain": { "eval": "'example.com'" }
        },
        "fields": {
            "name":   { "generator": "fake.company", "null_ratio": "0" },
            "email":  { "eval": "name.lower().replace(' ', '-') + '@' + email_domain" },
            "active": { "eval": "True" }
        }
    },
    {
        "operation": "write",
        "model": "res.partner",
        "ref": "my_partners",
        "batched": true,
        "values": {
            "suffix": { "eval": "'imported by populate'" }
        },
        "fields": {
            "comment": { "eval": "suffix" }
        }
    },
    {
        "operation": "function",
        "model": "res.partner",
        "name": "message_subscribe",
        "ref": "my_partners",
        "batched": true,
        "values": {
            "current_partner": { "eval": "env.user.partner_id.id" }
        },
        "args": {
            "0": { "eval": "[current_partner]" }
        }
    }
]

每个对象中可选的 fields、values 和 args 键把目标名称映射到其属性字典 —— 也就是你会作为 XML 属性写的那些键。位置参数使用数字字符串键("0"、"1" 等)。

字段、取值与参数

每个 <field/>、<value/> 和 <arg/> 声明都描述一个生成的目标:

  • <field/> 在 create 和 write 块中有效,并通过 ORM 持久化。

  • <value/> 在所有操作块中均有效,用于定义一个不持久化的局部变量。

  • <arg/> 在 function 块中有效,用于定义生成的方法参数。带有 name 的参数作为关键字参数传递;未命名的参数按声明顺序作为位置参数传递。

它们支持相同的生成属性,唯一区别是对于位置式 <arg/> 声明,name 是可选的。

Example

<field name="age" generator="scalar.integer" start="18" end="65"
       distribution="normal(mean=35, std=12)"/>
name (field 和 value 必填)

ORM 字段名、本地值名称或关键字参数名。

generator

要使用的生成器(参见 生成器)。与 eval 互斥。对于字段,如果两者都未提供,则会根据字段类型选择一个 默认生成器。值(values)和参数没有 ORM 类型,因此无法据此选择默认生成器。

eval

一个 Python 表达式。可通过名称引用其他生成目标以产生计算值。与 generator 互斥。

null_ratio

生成 False 而非真实值的概率(0–1)。默认值:0。不能与必填字段或带权重的 values 同时使用。

unique

设为 True 以强制唯一性。对于 ORM 字段,生成的值会与数据库中已有的记录以及同一作业(job)中先前生成的值进行检查。对于本地值和参数,由于它们不会被持久化,唯一性仅在当前作业范围内生效。

values

一个显式的值列表或带权重的字典。示例:"['a', 'b', 'c']" (权重相等)或 "{'a': 3, 'b': 1}" (a 被选中的可能性是 b 的 3 倍)。

distribution

一个统计分布规格说明,例如 "normal(mean=50, std=10)"。参见 分布。不能与带权重的 values 同时使用。

domain

一个 ORM 域(domain),用于过滤关联记录。仅适用于关系型(relational)和引用型(reference)生成器。可以包含在生成时解析的生成目标引用 – 参见 动态域(domain)。

ref

将关系型选取限定到在此引用标签下创建的记录。支持点路径(dot-path)遍历 – 参见 引用点路径导航。

comodel_name

对于关系型值或参数是必需的,即无法从 ORM 字段推断出 comodel 的情况。

partition

设为 True 以将 comodel ID 分区到并行的工作进程(workers)中。参见 并行执行的分区。

默认生成器

对于某个字段,如果既未指定 generator 也未指定 eval,则会根据字段类型自动选择一个默认生成器:

字段类型

默认生成器

boolean

scalar.boolean

integer

scalar.integer

float

scalar.float

monetary

scalar.monetary

char

textual.char

text

textual.text

html

textual.text

date

temporal.date

datetime

temporal.datetime

selection

choice.selection

binary

binary.binary

many2one

relation.one

one2many

relation.many

many2many

relation.many

many2one_reference

reference.one

reference

reference.raw

properties

properties.value

properties_definition

properties.definition

如果字段类型未在上表中列出,且既未提供 generator 也未提供 eval,则会引发错误。

生成器

生成器是生成字段、本地值和函数参数的基础构件。每个生成器都有一个 name (用于在蓝图 blueprint 中引用它)以及一组兼容的目标类型。

标量生成器

生成数值和布尔值。

scalar.boolean

生成 True 或 False。借助 values 可以为概率加权:values="{'True': 9, 'False': 1}" 约 90% 的时间会生成 True。

兼容目标:boolean 字段,以及生成的值或参数。

scalar.integer

在指定范围内生成随机整数。

兼容目标:integer 和 float 字段,以及生成的值或参数。

start

下界(含)。默认值:1。

end

上界(含)。默认值:1000000。

Example

<field name="quantity" generator="scalar.integer" start="1" end="100"/>
scalar.float

在指定范围内生成随机浮点数。

兼容目标:float 字段,以及生成的值或参数。

start

下界。默认值:1.0。

end

上界。默认值:1000000.0。

scalar.monetary

在指定范围内生成随机金额值。取决于模型(model)的货币字段 – 同一蓝图块中必须为该字段生成(或通过 eval 计算)一个值。

兼容目标:monetary 字段,以及生成的值或参数。

start

下界。默认值:1.0。

end

上界。默认值:1000000.0。

文本生成器

生成随机字符串。

textual.char

从字符集中生成固定长度的随机字符串。

兼容目标:char 和 html 字段,以及生成的值或参数。

char_set

候选字符集。默认值:ASCII 字母和数字。

length

生成字符串的长度。默认值:12。

textual.text

生成固定长度的随机文本块。

兼容目标:text 和 html 字段,以及生成的值或参数。

char_set

候选字符集。默认值:ASCII 字母、数字、空格和换行符。

length

生成文本的长度。默认值:50。

小技巧

如需逼真的文本(姓名、邮箱、地址),请改用 fake.* 生成器。

时间生成器

在指定范围内生成日期和日期时间,使用相对日期语法。

temporal.date

生成随机日期。

兼容目标:date 和 datetime 字段,以及生成的值或参数。

start

范围起点。默认值:None (时间之初)。

end

范围终点。默认值:None (时间之末)。

temporal.datetime

生成随机日期时间。

兼容目标:datetime 字段,以及生成的值或参数。

start

范围起点。默认值:None (时间之初)。

end

范围终点。默认值:None (时间之末)。

这两个生成器都支持在 start 和 end 中使用 相对日期语法 :

  • temporal.date 以 today 为锚点:"today -6m"、"today +1y"

  • temporal.datetime 以 now 为锚点:"now -30d"、"now +2h"

支持的单位后缀:y (年)、m (月)、w (周)、d (天)、h (小时)、M (分钟)、s (秒)。

Example

<field name="date_order" generator="temporal.date" start="today -6m" end="today"/>
<field name="create_date" generator="temporal.datetime" start="now -30d" end="now"/>

选择生成器

从一组值中选取。

choice.sample

从一个显式的 values 列表中选取(必填)。支持带权重的值。

兼容目标:integer、float、char、text、html、date、datetime、boolean 和 selection 字段,以及生成的值或参数。

Example

<field name="priority" generator="choice.sample"
       values="{'high': 1, 'medium': 5, 'low': 4}"/>
choice.selection

从字段自身的选择项键(selection keys)中选取。如果提供了 values,则仅使用其中的键(可带可选权重)。否则所有有效的选择项键被选中的可能性相等。

兼容类型:selection。

Example

<!-- All selection values equally likely -->
<field name="state" generator="choice.selection"/>

<!-- Only these values, with weights -->
<field name="state" generator="choice.selection"
       values="{'draft': 1, 'confirmed': 5, 'done': 3}"/>

二进制生成器

生成二进制数据。

binary.binary

生成随机二进制数据。

兼容目标:binary 字段,以及生成的值或参数。

size

大小(以字节计)。默认值:1024。

binary.image

生成一张随机纯色图像(PNG)。

兼容目标:binary 字段,以及生成的值或参数。

width

图像宽度(以像素计)。默认值:64。

height

图像高度(以像素计)。默认值:64。

关系型生成器

通过从已有记录中选取来生成关系型目标。

relation.one

选取一条关联记录。

兼容目标:many2one 字段,以及生成的值或参数。

domain

用于过滤候选记录的 ORM 域(domain)。参见 动态域(domain)。

ref

限定到此引用标签下创建的记录。参见 引用点路径导航。

comodel_name

对于生成的值或参数是必需的,即无法从 ORM 字段推断出 comodel 的情况。

partition

将 comodel ID 分区到并行的工作进程(workers)中。参见 并行执行的分区。

relation.many

选取多条关联记录(用于 one2many 和 many2many 字段)。

兼容目标:one2many 和 many2many 字段,以及生成的值或参数。

count

要建立关联的关联记录的平均数量。

std

数量的标准差。默认值:0 (始终恰好等于 count)。

groupby

按 comodel 上的某个字段对关联记录进行分组。

domain, ref, comodel_name, partition

与 relation.one 相同。

Example

<field name="tag_ids" generator="relation.many" count="3" std="2"/>

动态域(domain)

关系型生成器上的 domain 参数可以包含 生成目标引用 ,这些引用会在生成时针对当前记录已生成的值进行解析:

Example

<field name="project_id" generator="relation.one"/>
<field name="task_id" generator="relation.one"
       domain="[('project_id', '=', project_id)]"/>

域表达式中的 project_id 会被自动检测为依赖。生成时,表达式会使用 project_id 实际生成的值进行评估,因此每个 task_id 都能保证属于其兄弟字段 project_id。

引用点路径导航

ref 属性支持 点路径遍历 ,以将选取范围限定到先前创建批次中 相关 的记录上:

Example

<!-- Create projects and their tasks -->
<create model="project.project" count="10" id="my_projects">
    <field name="name" generator="fake.bs"/>
</create>
<create model="project.task" count="100">
    <field name="project_id" generator="relation.one" ref="my_projects"/>
</create>

<!-- Assign timesheets only to tasks that belong to our projects -->
<create model="account.analytic.line" count="200">
    <field name="task_id" generator="relation.one" ref="my_projects.task_ids"/>
</create>

ref="my_projects.task_ids" 通过获取 my_projects 下创建的记录、遍历 task_ids 关系,并将选取限定在这些 ID 上来解析。任何有效的对象关系映射点路径都可以使用。

这主要用于那些未在蓝图中显式创建的伴随记录,例如与 product.template 一起自动创建的 product.product 记录。

并行执行的分区

从伴随模型中选取的生成器(relation.one、relation.many、reference.one、reference.raw)支持 partition 参数。在并行作业中启用后,伴随模型 ID 会通过轮询分区在多个 worker 之间分配:

Example

<field name="user_id" generator="relation.one" partition="True"/>

这可以避免并行创建相关记录时发生冲突。

注解

  • 分区仅在作业存在兄弟子作业(即已被拆分以并行执行)时才生效。在单 worker 模式下,该参数不起作用。

  • 在与非均匀分布一起使用时,分区可能会引入轻微偏差。分布的整体形态得以保留,但参数不会那么精确地遵循。在大多数情况下可以忽略。

引用生成器

为引用类型字段生成值。

reference.one

为 many2one_reference 字段选取一条记录。隐式依赖于存储模型名的字段。

兼容类型:many2one_reference。

partition

在并行 worker 之间分区 ID。

reference.raw

为 reference 字段选取一条记录(存储 "model_name,id" 字符串)。

兼容类型:reference。

res_model

限定到特定模型。

res_id

限定到特定记录 ID。

ref

限定到此引用标签下的记录。

partition

在并行 worker 之间分区 ID。

Faker 生成器(fake.*)

封装了 Faker 库。任何受支持提供者的方法都可以直接使用为 fake.<method_name>:

Example

<field name="name"  generator="fake.name"/>
<field name="email" generator="fake.email" locale="fr_FR"/>
<field name="phone" generator="fake.phone_number"/>
<field name="bio"   generator="fake.paragraph" nb_sentences="5"/>

方法特定的关键字参数(例如 nb_sentences)会原样转发给 Faker 方法。

locale

用于本地化数据的区域设置。默认:en_US。

允许的提供者: address、automotive、bank、barcode、color、company、credit_card、currency、emoji、file、geo、internet、isbn、job、lorem、misc、passport、person、phone_number、profile、sbn、ssn、user_agent。

重要

Faker 必须单独安装。参见 安装。

其他生成器

misc.counter

生成等差数列。到达 end 时会回绕到 start。

兼容目标:integer 和 float 字段,以及生成的值或参数。

start

初始值。默认:0。

step

每条记录的增量。默认:1。

end

上界(会回绕)。默认:None (不回绕)。

Example

<field name="sequence" generator="misc.counter" start="1" step="1"/>
misc.cycle

按顺序确定性地循环遍历 values 列表。与 choice.sample 不同,这不是随机的——它会精确地重复该序列。

兼容目标:integer、float、char、text、html、date 和 datetime 字段,以及生成的值或参数。

注解

misc.cycle 不允许加权值——值总是按顺序循环。

Example

<field name="day" generator="misc.cycle"
       values="['Mon', 'Tue', 'Wed', 'Thu', 'Fri']"/>
misc.eval

计算一个 Python 表达式。可以引用其他目标名称来产生计算出的值。

兼容类型:任意。

求值上下文包含:

  • env —— Odoo 环境

  • model —— 正在填充数据的模型

  • Command —— odoo.fields.Command,用于构建关系命令

Example

<field name="display_name" generator="misc.eval"
       eval="name + ' (' + str(email) + ')'"/>

属性生成器

为 properties / properties_definition 字段系统生成值。

properties.definition

生成属性模式(属性定义列表)。

兼容类型:properties_definition。

props

显式的属性名称列表。

count

要生成的属性数量(在 props 未设置时使用)。

allowed_types

将生成的属性类型限定到该集合。

possible_values

针对 selection 类型的属性:将属性名称映射到其可能值的字典。

properties.prop

用于定义单个属性条目的辅助生成器。在 properties.definition 内部使用。

兼容目标:生成的值或参数。

prop_type

属性类型(例如 char、integer、selection)。

字符串

属性的显示标签。

possible_values

针对 selection 类型:可能值的列表。

properties.value

为 properties 字段生成值,与其父字段的 properties_definition 定义的模式相匹配。

兼容类型:properties。

分布

默认情况下,生成器在其取值范围内均匀随机地产生值。添加 distribution 参数会改变 哪些部分 被采样的 可能性 。

Example

<field name="age"   generator="scalar.integer" start="18" end="90"
       distribution="normal(mean=35, std=12)"/>
<field name="delay" generator="scalar.float"   start="0"  end="100"
       distribution="exponential(rate=0.05)"/>

normal(mean, std) —— 大部分值集中在中间

产生经典的钟形曲线。大部分值落在 mean 附近;距离越远,越少见。std (标准差)控制分布的离散程度——std 越小,值越紧密地聚集在均值周围。

适用场景 :你想要真实的”均值带自然波动”模式。

示例字段

参数

原因

员工年龄

normal(mean=35, std=12)

大多数员工在 35 岁左右,很年轻或很年长的较少

产品价格

normal(mean=50, std=15)

价格集中在 50 附近,有一些偏低或偏高的异常值

任务时长(小时)

normal(mean=8, std=3)

大多数任务约需一天,有些更短或更长

uniform() —— 任何值被取到的概率相同

一种平坦的分布——取值范围内的每个值被选中的机会完全相同。当你完全省略 distribution 时,这就是默认行为,所以你很少需要显式写出它。

exponential(rate) —— 大量小值,少量大值

一条起始较高然后衰减的陡峭曲线。大部分生成的值较小;较大的值越来越罕见。rate 越高,衰减越快。

适用场景 :数据应偏向低端,并偶尔出现尖峰。

示例字段

参数

原因

距截止日的天数

exponential(rate=0.03)

大多数截止日期临近,少数在几个月后

已分配小时数

exponential(rate=0.1)

大多数任务较快,少数很长

事件间隔时间

exponential(rate=0.05)

短间隔常见,长间隔罕见

beta(alpha, beta) —— 0 到 1 之间的值,形状由你决定

总是产生 [0, 1] 内的值。生成器会自动将其映射到 start/end 范围。这两个参数决定了曲线的形状:

  • alpha=2, beta=2 —— 钟形,居中于 0.5(类似有界正态分布)

  • alpha=1, beta=3 —— 偏向 0(大部分值较低)

  • alpha=3, beta=1 —— 偏向 1(大部分值较高)

  • alpha=0.5, beta=0.5 —— U 形,值聚集在 0 和 1 附近

适用场景 :你正在建模百分比、进度、评分或任何有界比例。

示例字段

参数

原因

项目进度(%)

beta(alpha=2, beta=2)

大多数项目大约进行到一半,0% 或 100% 的很少

折扣率

beta(alpha=1, beta=3)

大多数折扣较小,大折扣罕见

满意度评分

beta(alpha=3, beta=1)

大多数评分偏高

poisson(lam) —— 某件事发生多少次

产生整数,表示 发生次数 。lam (lambda)是你期望的平均发生次数。接近 lam 的值最可能出现;远离它的值罕见。

适用场景 :你在生成”多少”——例如物品、事件或尝试的次数。

示例字段

参数

原因

订单行数

poisson(lam=5)

订单平均 5 行,有些只有 1 行,15 行以上的很少

每日支持工单数

poisson(lam=3)

平均每天约 3 个

登录尝试次数

poisson(lam=2)

通常是 1–3 次尝试,偶尔更多

triangular(min, max, mode) —— 三点估算

一个简单的三角形形状。mode 是峰值(最可能的值),min 和 max 是绝对边界。接近 mode 的值最常见;概率向两端线性衰减。

适用场景 :你能估算三个点——最小值、最大值和最可能值——但没有更详细的数据。

示例字段

参数

原因

任务预估(天)

triangular(min=1, max=30, mode=5)

大多数任务约 5 天,从不会少于 1 天或多于 30 天

运费

triangular(min=5, max=200, mode=25)

通常在 25 左右,边界为 5 和 200

快速决策指南

你想要……

使用

围绕均值的真实聚集

normal

任何值被取到的概率相同

uniform (或省略 distribution)

多数是小值,大值罕见

exponential

百分比/有界比例

beta

“多少次”的计数

poisson

三点估算(最小值/最可能值/最大值)

triangular

高级主题

生成的值

字段是传递给 ORM create 或 write 的持久化值。相比之下,<value/> 声明是生成的但 不持久化 的局部变量。它们允许你构建多个字段依赖的中间值,从而避免重复:

Example

<create model="account.move.line" count="1000">
    <field name="quantity"    generator="scalar.integer" start="1" end="100"/>
    <field name="price_unit"  generator="scalar.float"   start="5" end="500"/>
    <value name="subtotal" eval="quantity * price_unit"/>
    <field name="discount" eval="subtotal * 0.1 if subtotal > 200 else 0"/>
    <field name="price_total" eval="subtotal - discount"/>
</create>

此处 subtotal 被计算但从不写入数据库。discount 和 price_total 都引用它,因此 quantity * price_unit 的逻辑只在一处存在。

生成的值同样可用于 关联 多个持久化字段:

Example

<create model="res.partner" count="200">
    <value name="first_name" generator="fake.first_name"/>
    <value name="last_name"  generator="fake.last_name"/>
    <field name="name"  eval="first_name + ' ' + last_name"/>
    <field name="email"
           eval="first_name.lower() + '.' + last_name.lower() + '@example.com'"/>
</create>

每条记录的 name 和 email 彼此保持一致,而两个中间值本身都不被单独存储。

写入任务

使用 <write/> 更新已有记录。目标记录可以通过 ref、顶层 domain 或两者组合来选取。

Example

<!-- Create partners and tag them with the "customers" reference -->
<create model="res.partner" count="500" id="customers">
    <field name="name" generator="fake.company"/>
    <field name="active" values="{'True': 9, 'False': 1}"/>
    <field name="customer_rank" start="0" end="5"/>
</create>

<!-- Update all partners created under the "customers" reference -->
<write model="res.partner" ref="customers">
    <field name="phone" generator="fake.phone_number"/>
</write>

<!-- Update all active customers, even if they were not created by this blueprint -->
<write model="res.partner"
       domain="[('customer_rank', '&gt;', 0), ('active', '=', True)]">
    <field name="mobile" generator="fake.phone_number"/>
</write>

<!-- Update only active customers created under the "customers" reference -->
<write model="res.partner" ref="customers"
       domain="[('customer_rank', '&gt;', 0), ('active', '=', True)]">
    <field name="email" generator="fake.company_email"/>
</write>

选取规则如下:

属性

目标记录

仅 ref

在该 populate 引用下创建的记录。

仅 domain

作业模型中符合 domain 条件的既有记录。

ref 和 domain

既被引用又符合 domain 条件的记录。

两者皆非

作业模型中所有的既有记录。

write 作业上的 domain 只会评估一次以选取目标记录,并非按每条生成记录动态计算。create 作业不能定义顶层 domain,因为它们创建新记录,而不是针对既有记录。

默认情况下,每条目标记录生成并写入一次值。使用 batched="True" 时,整个可执行作业或子作业只生成一组值,并通过一次 ORM write 调用更新该记录集:

Example

<write model="res.partner" ref="customers" batched="True">
    <field name="active" eval="True"/>
</write>

batched 仅在 write 和 function 块上有效。create 块不接受该属性,因为 ORM create 已经接收一组生成值的列表。大型作业仍然可以拆分为子作业,每个子作业执行自己的写入。

函数任务

使用 <function/> 在目标记录上调用模型方法。当记录无法直接以最终业务状态创建、必须经过一个过渡方法时,这非常有用。例如,发票先以草稿状态创建,再通过调用 _post 将其过账:

Example

<function model="account.move" name="_post" ref="moves" batched="True">
    <arg name="soft" eval="False"/>
</function>

函数块遵循与写入块相同的 ref 和 domain 选取规则。name 属性选择要调用的方法。带有 @api.model 装饰器的方法会在空模型记录集上每个作业调用一次;常规记录方法则在目标记录上调用。

使用 <arg/> 声明方法参数。参数支持与 <value/> 相同的生成属性。命名参数成为关键字参数,无命名参数则按声明顺序成为位置参数:

Example

<function model="x.model" name="action" ref="records">
    <arg eval="'first positional'"/>
    <arg eval="42"/>
    <arg name="flag" eval="True"/>
</function>

在 JSON 中,位置参数使用数字字符串键,如 "0" 和 "1"。默认 batched="False" 时,参数会被生成且方法每条目标记录调用一次。batched="True" 时,只生成一组参数,并在每个可执行作业或子作业的目标记录集上调用一次方法。这一批处理差异仅适用于常规记录方法;@api.model 方法总是每个作业调用一次。

蓝图继承

蓝图通过 inherit_id 支持 Odoo 风格的视图继承。子蓝图会对父蓝图的 XML 定义应用 XPath 或位置规格:

Example

<record id="custom_blueprint" model="populate.blueprint">
    <field name="name">Custom Blueprint</field>
    <field name="inherit_id" ref="base_module.parent_blueprint"/>
    <field name="definition_xml" type="xml">
        <!-- Change record count -->
        <create model="res.partner" position="attributes">
            <attribute name="count">2000</attribute>
        </create>
        <!-- Add a new field to an existing create block -->
        <create model="res.partner" position="inside">
            <field name="website" generator="fake.url"/>
        </create>
        <!-- Add a new create block after an existing one -->
        <create model="res.partner" position="after">
            <create model="res.users" count="50" id="new_users">
                <field name="name" generator="fake.name"/>
                <field name="login" generator="fake.user_name" unique="True"/>
            </create>
        </create>
    </field>
</record>

支持的位置:attributes、inside、before、after、replace。XPath 表达式(<xpath expr="..." position="...">)同样有效。链式继承(孙级蓝图)受支持;循环继承会被检测并拒绝。

导入蓝图

蓝图继承 扩展单个蓝图,而 <import/> 将多个独立的 XML 蓝图组合成新的场景。导入的操作块会插入到导入所在的位置,并成为调用方会话的普通作业:

Example

<create model="sale.order" count="500" id="orders">
    <!-- ... -->
</create>

<import ref="product.fake_product_demo" as="catalog"/>

<create model="sale.order.line" count="2000">
    <field name="order_id" ref="orders"/>
    <field name="product_id"
           ref="catalog/product_templates.product_variant_ids"/>
</create>

ref 属性必须包含另一个 populate.blueprint 记录的全限定外部 ID。<import/> 必须是蓝图 <data/> 根元素的直接子元素。

可选的 as 属性会为导入蓝图中声明的每个 id 以及其内部引用添加命名空间。这可以避免同一场景导入多个具有相同 ID 的蓝图,或多次导入同一蓝图时产生冲突。若不使用 as,ID 将保持不变,且不得与调用方或其他导入中的 ID 冲突。

斜杠(/)用于分隔命名空间组件,而点号(.)保留其 遍历关系型字段 的既有含义。例如在 catalog/product_templates.product_variant_ids 中,catalog/product_templates 是 populate 引用,product_variant_ids 是关系型字段路径。

通过把 inherit_id 支持的相同继承规格作为 <import/> 的子元素添加,即可在不修改源蓝图的情况下定制导入:

Example

<import ref="product.fake_product_demo" as="catalog">
    <xpath expr="//create[@id='product_templates']" position="attributes">
        <attribute name="count">250</attribute>
    </xpath>
    <xpath expr="//create[@id='product_supplier_info']" position="replace"/>
</import>

XPath 表达式针对源蓝图的原始 ID,因为规格在命名空间之前应用。对组合后的蓝图应用的继承则看到的是加了命名空间的 ID,例如 catalog/product_templates。

导入是递归解析的,因此其命名空间会组合。例如,在 sales 命名空间下导入 catalog 命名空间会产生 sales/catalog/products 之类的引用。导入与继承的递归组合会被拒绝。

重要

仅 XML 蓝图支持导入;只有 JSON 的蓝图无法被导入。调用方模块必须依赖提供被导入蓝图的模块。当两者位于同一模块时,被导入的蓝图必须在该模块的 populate 数据文件中更早定义。

导入在会话创建时解析。新会话使用最新的蓝图定义,而已存在的会话则保留其创建时可用定义所实例化的作业。

会话与恢复

每次运行都会创建一个 Session (populate.session),用于跟踪每个作业及其产生的记录。如果执行被中断(Ctrl+C、崩溃等),你可以从上次中断处继续:

Example

# Resume the most recent unfinished session
$ odoo-bin populate -d mydb --resume

# Resume a specific session by ID
$ odoo-bin populate -d mydb --resume 42

会话还保证 确定性生成 :使用相同的 --seed 和同一蓝图,每次都会产生相同的数据。

并行执行

传入 -j (或 -j auto)可将大型作业拆分到多个工作进程。每个超过内部批处理大小的作业都会被自动划分为子作业并分发到进程池。

Example

$ odoo-bin populate -d mydb -b my_blueprint --scale 50 -j auto

当模型的约束要求顺序写入时,可以通过 parallel="False" 逐个操作块禁用并行。多进程后端由环境变量 ODOO_POPULATE_MULTIPROCESS_ENABLE 控制(默认为 True)。

约束违规的自动重试

会话执行器包含针对瞬时数据库约束失败的重试机制。当作业触发以下任一 PostgreSQL 违规时,该作业的 seed 会被重新掷出,整个作业会带着一组全新的随机值重新执行(最多尝试 5 次):

违规

常见原因

提示

UniqueViolation

两条生成记录在唯一索引上冲突

使用能产生更多样化值的生成器,或添加 unique="True"

NotNullViolation

必填列收到 NULL

在该字段上添加 null_ratio="0"

CheckViolation

某个生成值未通过 CHECK 约束

调整生成器参数使其落在约束范围内

ExclusionViolation

生成值违反了排斥约束

调整生成器参数使其落在约束范围内

这意味着蓝图无需在前期完美调优——因随机性导致的偶发约束失败会被透明处理。只有在所有重试后仍然存在的违规才会以错误形式浮出。

编写自定义生成器

你可以通过继承 odoo.addons.populate.generators.Generator 创建自定义生成器,并将代码放入模块的 populate/ 包(需带 __init__.py)。模块加载时生成器会自动注册。

Example

from odoo.addons.populate.generators import Generator


class SequentialEmail(Generator):
    """Generates email addresses like user_0001@example.com, user_0002@example.com, ..."""

    name = 'my_module.sequential_email'
    allowed_on = ('char', 'value')

    def __init__(self, domain_name='example.com', **kwargs):
        super().__init__(**kwargs)
        self.domain_name = domain_name
        self._counter = 0

    def _next(self, known_vals):
        self._counter += 1
        return f'user_{self._counter:04d}@{self.domain_name}'

    @classmethod
    def convert_to_kwargs(cls, attrs):
        kwargs = super().convert_to_kwargs(attrs)
        if 'domain_name' in attrs:
            kwargs['domain_name'] = attrs['domain_name']
        return kwargs

关键要求:

name (类属性,必需)

生成器的唯一字符串标识符。约定:<module_name>.<generator_name>。

allowed_on (类属性,可选)

兼容的 字段类型 元组。包含 value 允许生成器用于生成值和参数。设为 None 则允许任意目标。

_next(self, known_vals) (方法,必需)

生成并返回下一个值。known_vals 是当前记录目标名到其已生成值的字典(只有列入 depends 的目标能保证存在)。

convert_to_kwargs(cls, attrs) (类方法,可选)

覆盖该方法,将 XML/JSON 属性转换为 __init__ 的关键字参数。务必先调用 super().convert_to_kwargs(attrs) 以处理标准属性(values、null_ratio、distribution、unique)。

注册之后,该生成器即可在任何蓝图中使用:

Example

<field name="email" generator="my_module.sequential_email"
       domain_name="mycompany.com"/>

指南

以下准则并非硬性规定,但有助于你编写蓝图并处理一些边界情况。

选择可跨层级平滑扩展的数量

蓝图的 count 值应在 --scale 1 时产生一个实用、可浏览的数据集,并在更高比例下保持连贯。一种可行的做法是对齐三个层级:

  • 1x(基础)——独立的演示或开发数据集。体量足够验证分页、搜索和筛选,又足够小到可以在几秒内填充完毕。主交易模型约 10 000 条记录是合理的基线。

  • 10x——负载测试规模(约 100 000 条记录)。揭示 UI 卡顿和未建索引的查询瓶颈。

  • 100x——压力测试规模(约 1 000 000 条记录)。暴露对象关系映射或 PostgreSQL 的可扩展性上限。

Example

$ odoo-bin populate -d mydb -b my_module.demo              # 1x  — 10 000 tasks
$ odoo-bin populate -d mydb -b my_module.demo --scale 10   # 10x — 100 000 tasks
$ odoo-bin populate -d mydb -b my_module.demo --scale 100  # 100x — 1 000 000 tasks

保持比例合理。 绝对数量不如相关模型之间的比例重要。如果你创建 120 个项目和 10 000 个任务,那大约是每项目 80 个任务——一个合理的平均值。在 100x 下,这变成 12 000 个项目和 1 000 000 个任务,比例保持不变。

主数据不受 scale 影响。 阶段定义、属性集及类似的配置记录应始终使用 scale="False",以便在所有层级保持固定数量。1x 下 8 个任务阶段,100x 下仍然是 8 个任务阶段。

使用 context 禁用副作用

禁用邮件通知、字段跟踪和自动记录创建后,批量填充会显著更快。

Example

<create model="account.move" count="15000" id="invoices"
        context="{'mail_auto_subscribe_no_notify': True}">
    ...
</create>

使用 partition="True" 避免多工作进程模式下的序列化错误

在并行创建子记录时(例如为订单创建订单行),只要父模型存在 依赖子记录的存储型计算字段 ,就应该在父字段上添加 partition="True"。

Example

<create model="sale.order.line" count="20000" id="order_lines">
    <field name="order_id" ref="sale_orders" partition="True"/>
</create>

不进行分区时,工作进程会随机挑选父 ID。两个工作进程可能同时为同一订单创建行。由于 sale.order 有在 order_line 变化时重算的存储型计算字段,两个工作进程会同时尝试写入同一订单行。PostgreSQL 检测到这一冲突并抛出 序列化错误 。

使用 partition="True" 后,每个工作进程被分配到互不重叠的父 ID 子集。任意两个工作进程都不会触及同一父记录,因此并发写入不会冲突,也不会发生序列化错误。

用生成值承载中间逻辑

生成值不会写入数据库,但能让蓝图更清晰、更易于维护。适用场景:

关联字段——生成一次值,在多个持久化字段中复用:

Example

<value name="first_name" generator="fake.first_name"/>
<value name="last_name"  generator="fake.last_name"/>
<field name="name" eval="first_name + ' ' + last_name"/>
<field name="email"
       eval="first_name.lower() + '.' + last_name.lower() + '@example.com'"/>

多字段唯一性——把多个字段打包成元组并标记为唯一,然后再解包:

Example

<value name="generated_product_id" generator="relation.one"
       comodel_name="product.product" ref="products"/>
<value name="generated_partner_id" generator="relation.one"
       comodel_name="res.partner" ref="customers"/>
<value name="unique_pair" eval="(generated_product_id, generated_partner_id)" unique="True"/>
<field name="product_id" eval="unique_pair[0]"/>
<field name="partner_id" eval="unique_pair[1]"/>

当两个字段之间存在复合唯一约束,而只在其中一个字段上加 unique=True 又会过度限制可能的组合时,此项必不可少。

计算数量——推导出一个比例,再将其应用:

Example

<value name="ratio" generator="scalar.float"
       start="0" end="1" distribution="beta(alpha=2, beta=2)"/>
<field name="qty_delivered" eval="product_uom_qty * ratio"/>

使用 eval 从父记录派生值

当子记录需要一个与父记录匹配的值(例如子任务继承其父任务的项目)时,使用配合 model.browse() 或 env[...] 的 eval:

Example

<!-- Subtasks inherit the project from their parent task -->
<field name="parent_id" ref="parent_tasks"/>
<field name="project_id" eval="model.browse(parent_id).project_id.id"/>

<!-- Invoice currency matches the journal's currency -->
<field name="currency_id"
       eval="(journal := env['account.journal'].browse(journal_id)).currency_id.id
             or journal.company_id.currency_id.id"/>

使用写入块实现两阶段创建

某些模型要求按特定顺序设置字段,或需要二次处理以模拟真实的状态转换。使用 <write/> 更新先前创建的记录:

Example

<!-- Phase 1: create product templates without variants -->
<create model="product.template" count="5000" id="templates"
        context="{'create_product_product': False}">
    <field name="name" generator="fake.catch_phrase"/>
</create>

<!-- Phase 2: add attribute lines (triggers variant creation) -->
<create model="product.template.attribute.line" count="8000" id="attr_lines">
    <field name="product_tmpl_id" ref="templates"/>
    ...
</create>

<!-- Phase 3: update the generated variants -->
<write model="product.product" ref="templates.product_variant_ids">
    <field name="default_code" generator="fake.ean13" unique="True"/>
</write>