Skip to content

Camera#

The camera model changes an ideal pinhole render into the image of a real machine-vision camera. This page gives the capture procedure, the lens equations, the rolling shutter and the image sensor chain.

The camera pipeline: the scene capture renders a wide pinhole image, the GPU readback copies it to the CPU, the camera model applies the lens, the rolling shutter and the image sensor, and the result goes to a PNG file, to the sensor stream and to the HUD. RENDER, GAME THREAD AND GPU Scene capture pinhole render, wider than the output FinalColorLDR, 8-bit sRGB Warm-up and sample renders N − 1 renders before the due time, then the sample render GPU readback ring of 3 readback objects, copy to the CPU without a stall Camera model thread pool, for each output pixel render pixels and camera motion 1. Undistorted ray Newton inverse of Brown-Conrady, one table for all frames 2. Chromatic aberration red ray × (1 + c), blue ray × (1 − c) 3. Rolling shutter row time, then rotation and translation of the ray 4. Projection and sample pinhole projection into the render, bilinear sample in linear light 5. Vignetting gain from the radius, poly or cos4 6. Bayer site and electrons one colour for each pixel, PRNU, shot, dark and read noise 7. ADC black level, round and clip, 12 bits 8. Demosaic Malvar-He-Cutler 5 × 5 filters or bilinear 9. sRGB encode 8 bits for each channel, image of the output size PNG file and camera row camera/front_<time>.png Sensor stream frame camera frame, bgr8 HUD view half resolution
The camera pipeline: the render on the GPU, the nine stages of the camera model and the three outputs. Open the diagram

Scope and Assumptions#

The game has one front camera for each vehicle. ACRES Core has no camera.

The model has two parts.

  1. Unreal Engine renders an ideal pinhole image that is wider than the output image.
  2. FAcresCameraModel::Process makes each output pixel from that render on the CPU.

The model makes these assumptions.

  • The render is an 8-bit image after exposure and tone mapping (FinalColorLDR). The model applies the inverse sRGB curve and uses the result as a linear sensor signal. A value of 1.0 fills the pixel at the base ISO.
  • Noise thus scales with the brightness of the render, not with the radiance of the scene.
  • One render gives one image. The rolling shutter is a warp of that render with a constant camera motion and one scene depth.
  • The camera has no white balance, no colour correction matrix, no noise reduction and no sharpening.
  • The exposure time adds no motion blur.

The INS, the GNSS receiver and the CAN signals are on other pages: INS and GNSS and Wheel and CAN Signals. The procedures to record the sensors are in Sensors and Rigs.

Symbols#

Symbol Quantity Unit
\(W, H\) Width and height of the output image px
\(u, v\) Column and row of an output pixel. Pixel centres are at integers. px
\(f_x, f_y\) Focal lengths of the output image px
\(c_x, c_y\) Principal point of the output image px
\(x, y\) Undistorted normalised ray coordinates, \(x = X/Z\), \(y = Y/Z\) -
\(x_d, y_d\) Distorted normalised coordinates -
\(r, r_d\) Undistorted and distorted radius -
\(k_1, k_2, k_3\) Radial distortion coefficients -
\(p_1, p_2\) Tangential distortion coefficients -
\(s_{min}, r_c\) Minimum radial slope and the radius where the linear extension starts -
\(c\) Ray scale of the chromatic aberration -
\(W_r, H_r, f_r, c_{xr}, c_{yr}\) Size, focal length and principal point of the render px
\(T_{ro}\) Readout time of the rolling shutter, first row to last row s
\(\Delta t_v\) Time of row \(v\) relative to the render s
\(\vec{\omega}, \mathbf{v}_c\) Angular velocity and linear velocity of the camera in optical axes rad/s, m/s
\(Z_a\) Assumed scene depth m
\(L\) Linear signal of a pixel, 0 to 1 -
\(V\) Vignetting gain -
\(N, t_{exp}, I\) F-number, exposure time and ISO -, s, -
\(g\) Analogue gain, \(I / I_{base}\) -
\(E_{fw}\) Full-well capacity e⁻
\(\sigma_{read}\) Read noise e⁻
\(I_{dark}\) Dark current e⁻/s
\(\sigma_{prnu}\) Standard deviation of the pixel gain -
\(b, B\) ADC bits and black level -, DN
\(n\) A standard normal number -

The optical frame has x to the right, y down and z forward. The body frame FLU has x forward, y to the left and z up.

Capture#

Source: FAcresSensorRecorder::Start and FAcresSensorRecorder::TickCamera in AcresSensors.cpp.

Scene Capture#

A scene capture component renders the scene into a render target of the format RGBA8 sRGB. The component attaches to the chassis at camera.position_flu_m with the rotation camera.rpy_deg. The rotation uses the ROS convention: roll, pitch and yaw about the fixed FLU axes, \(R = R_z R_y R_x\). A positive pitch points the camera down.

The capture source is FinalColorLDR: the image after tone mapping, as the main view shows it. The capture has these settings.

  • Motion blur is off.
  • Lumen global illumination and Lumen reflections are on, with hardware ray tracing when the project has it.
  • Temporal anti-aliasing is on when camera.temporal_history is true.
  • The capture keeps its render state between renders, thus the temporal history stays valid.

The render state of the capture follows the sensor profile, not the render tier. Refer to Platforms and GPU Tiers.

Exposure#

The key camera.exposure.mode selects the exposure of the capture.

Mode Menu Name Exposure
volume Auto (Scene) The post-process volume of the level sets the exposure. This is the default.
physical Manual A fixed exposure from the f-number, the exposure time and the ISO.
copy - The capture copies the settings of the unbound post-process volume with the highest priority.

In the mode physical the capture uses manual exposure with the physical camera values. The exposure value at ISO 100 is

\[ EV_{100} = \log_2\!\left(\frac{N^2}{t_{exp}} \cdot \frac{100}{I}\right) \]

The shipped values \(N = 8\), \(t_{exp} = 0.003\) s and \(I = 100\) give \(EV_{100} = 14.4\). The exposure compensation is 0.

In all modes the image sensor model uses two of these values. The ISO sets the analogue gain \(g\). The exposure time sets the dark signal \(I_{dark}\, t_{exp}\).

Render Size#

Barrel distortion makes the output image see a wider field than a pinhole image of the same focal length. The rolling shutter also moves the rows. The render is thus larger than the output. FAcresCameraModel::Initialize computes the render size from the largest undistorted ray coordinates \(X_{max}\) and \(Y_{max}\) of all output pixels.

\[ f_r = f_x\, s_r, \qquad G = (1 + c)(1 + m) \]
\[ W_r = 2 \left\lceil f_r X_{max} G + 1 \right\rceil, \qquad H_r = 2 \left\lceil f_r Y_{max} G + 1 \right\rceil \]
\[ c_{xr} = \frac{W_r - 1}{2}, \qquad c_{yr} = \frac{H_r - 1}{2}, \qquad \mathit{HFOV}_r = 2 \arctan\frac{W_r}{2 f_r} \]

\(s_r\) is lens.render_scale and \(m\) is lens.render_margin. The game stops when \(W_r\) or \(H_r\) is larger than 8192. With the model off, the render has the size and the field of view of the output.

Camera Output Render Render HFOV Effective Output FOV
Default (sensors.json) 896 x 512 1024 x 586 97.63° 94.5° x 60.7°
Polaris (sensors_polaris.json) 896 x 512 1316 x 770 114.94° 100.07° x 51.61°

The effective field of view uses the row and the column through the principal point. It is the angle between the rays of the first pixel and the last pixel.

Warm-Up Frames and Capture Scheduling#

Lumen, temporal super-resolution and the volumetric clouds collect data from many frames. A capture that renders one time has no such history: the shadows are black and the clouds smear. The recorder thus renders the capture on some frames before each sample.

The sensor pipeline: the physics step is the clock, a due-time scheduler starts the samples, the physics thread, the game thread and the thread pool make the data, and the data go to the session log and to the sensor stream. CLOCK AND SAMPLE SCHEDULING Physics step clock of all sensors, 1/120 s Due-time scheduler first sample at t = 0, then one sample each 1/f, in the first step at or after the due time truth, actions, IMU 120 Hz · INS 100 Hz · CAN 20 Hz · GNSS, LiDAR, camera 10 Hz samples that are due camera due time SENSOR MODELS, BY THREAD Physics thread truth, actions and CAN rows IMU, GNSS and INS models pose history of the LiDAR sweep request for each LiDAR scan Game thread LiDAR: beam trace, GPU ray pass, echoes, detection and returns camera: warm-up renders, sample render, GPU readback Thread pool LiDAR: organised cloud, PLY and PCD encode camera: camera model, PNG encode LATENCY AND FILES Latency buffer available = sample + delay_s release in a physics step Writer thread one thread for all rows queue limit 200 000 lines Session log JSONL rows, episode.json PLY, PCD and PNG files The LiDAR row and the camera row go into the latency buffer when their files exist. The truth, actions and CAN rows go to the writer thread without a delay. rows PLY, PCD, PNG clock, INS cloud, image Sensor stream sender thread binary frames on 127.0.0.1, option -SensorStream=, no modelled delay, queue limit 96 MiB for clouds and images ROS 2 bridge sim_bridge publishes the ROS 2 topics
The sensor pipeline: the physics step is the clock of all sensors. The data go to the session log and to the sensor stream. Open the diagram

TickCamera runs one time in each rendered frame on the game thread. \(t\) is the time of the latest physics step, \(t_{next}\) is the next due time and \(P = 1/f_{cam}\) is the sample period. \(\bar{t}_f\) is the smoothed frame time: \(\bar{t}_f \leftarrow 0.9\,\bar{t}_f + 0.1 \min(\Delta t, 0.25)\). \(N_w\) is camera.warmup_frames.

  1. Before the due time, the capture renders a warm-up frame when \(t + (N_w - 0.5)\,\bar{t}_f \ge t_{next}\). It renders \(N_w - 1\) warm-up frames at most.
  2. At the due time, a hitch can leave fewer than \(N_w - 1\) warm-up frames. The capture then renders one more warm-up frame and the sample moves to the next frame.
  3. The recorder counts the samples that the frame missed: \(M = \lfloor (t - t_{next}) / P \rfloor\). Then \(t_{next} \leftarrow t_{next} + (M + 1) P\).
  4. The recorder makes a capture job with the time \(t\), the camera pose and the camera motion.
  5. The capture renders the sample at the end of this frame.
  6. In the next frame the recorder starts the GPU readback. It examines the readback in each subsequent frame.
  7. When the pixels are on the CPU, a thread-pool task applies the camera model and writes the outputs.

The first sample is at \(t = 0\). The sample time is the time of a physics step, but the step is not always the due step.

Setting Capture Method
temporal_history true, warmup_frames > 0 Warm-up frames, then the sample render. This is the default (2 frames).
temporal_history true, warmup_frames = 0 The capture renders in each frame.
temporal_history false One render on demand for each sample, with no history.

Readback#

Three readback objects make a ring. One sample uses one object until its pixels are on the CPU. The readback of one image can thus overlap the warm-up of the next image.

The recorder drops a sample and increments a counter in these conditions.

Counter in episode.json Condition
camera_missed A frame came more than one sample period late.
camera_busy Three samples are in the readback ring.
camera_encode_backlog Four encode tasks of the camera and the LiDAR run at the same time.

Intrinsics#

Source: FAcresCameraModel::Initialize in AcresCameraModel.cpp.

The model uses the OpenCV pinhole convention. The centre of the top left pixel is \((0, 0)\).

\[ u = f_x\, x_d + c_x, \qquad v = f_y\, y_d + c_y \]

Without the key lens.intrinsics, the horizontal field of view camera.hfov_deg gives the focal length.

\[ f_x = f_y = \frac{W}{2 \tan(\mathit{HFOV}/2)}, \qquad c_x = \frac{W-1}{2} + o_x, \qquad c_y = \frac{H-1}{2} + o_y \]

\((o_x, o_y)\) is lens.principal_offset_px. The model adds it only when the lens model is on. The shipped values give \(f_x = f_y = 448\) px, \(c_x = 451.0\) px and \(c_y = 253.0\) px.

With the key lens.intrinsics, the model uses the calibrated \(f_x, f_y, c_x, c_y\) directly and ignores the offset. A change of the output size keeps the field of view of a calibrated camera. For a new width \(W'\) the scale is \(s = W'/W\):

\[ f_x' = s\, f_x, \qquad c_x' = (c_x + 0.5)\, s - 0.5 \]

The same equations apply to \(f_y\) and \(c_y\) with the heights.

Lens Distortion#

Source: Distort and FAcresCameraModel::Initialize in AcresCameraModel.cpp.

The lens uses the Brown-Conrady model in the form of the OpenCV model plumb_bob. With \(r^2 = x^2 + y^2\) and the radial factor \(\rho(r) = 1 + k_1 r^2 + k_2 r^4 + k_3 r^6\):

\[ \begin{aligned} x_d &= x\,\rho + 2 p_1 x y + p_2 (r^2 + 2x^2) \\ y_d &= y\,\rho + p_1 (r^2 + 2y^2) + 2 p_2 x y \end{aligned} \]

Linear Extension#

A calibration with strong barrel distortion can turn back before the corners of the image. The distorted radius \(r_d = r\,\rho(r)\) then decreases and no ray reaches the corners. When lens.radial_min_slope \(= s_{min} > 0\), the model limits the slope of \(r_d(r)\).

The limit radius \(r_c\) is the first radius in the sequence 0.001, 0.002, ... 3.000 where

\[ \frac{d r_d}{d r} = 1 + 3 k_1 r^2 + 5 k_2 r^4 + 7 k_3 r^6 \le s_{min} \]

For \(r > r_c\) the radius continues as a straight line with that slope:

\[ r_d = A + s_{min}(r - r_c), \qquad A = r_c\,\rho(r_c), \qquad \rho(r) = s_{min} + \frac{A - s_{min} r_c}{r} \]

Inverse#

The model needs the ray \((x, y)\) of each output pixel. The inverse has no closed form. Initialize solves it with the Newton method for each pixel and stores the result in a table.

\[ x_d^{*} = \frac{u - c_x}{f_x}, \qquad y_d^{*} = \frac{v - c_y}{f_y}, \qquad \mathbf{e} = \begin{pmatrix} x_d(x, y) - x_d^{*} \\ y_d(x, y) - y_d^{*} \end{pmatrix} \]
\[ \begin{pmatrix} x \\ y \end{pmatrix} \leftarrow \begin{pmatrix} x \\ y \end{pmatrix} - J^{-1} \mathbf{e} \]

The Jacobian is symmetric. With \(D = d\rho / d(r^2)\):

\[ J = \begin{pmatrix} \rho + 2x^2 D + 2 p_1 y + 6 p_2 x & 2xyD + 2 p_1 x + 2 p_2 y \\ 2xyD + 2 p_1 x + 2 p_2 y & \rho + 2y^2 D + 6 p_1 y + 2 p_2 x \end{pmatrix} \]
\[ D = k_1 + 2 k_2 r^2 + 3 k_3 r^4 \quad (r \le r_c), \qquad D = -\frac{A - s_{min} r_c}{2 r^3} \quad (r > r_c) \]

The start value is \((x, y) = (x_d^{*}, y_d^{*})\). The iteration stops when the two components of \(\mathbf{e}\) are smaller than \(10^{-12}\), or after 30 steps.

The game stops with a message in two conditions.

  • The largest residual \(\max\lvert\mathbf{e}\rvert f_x\) is 0.001 px or more.
  • The slope \(1 + 3 k_1 r^2 + 5 k_2 r^4 + 7 k_3 r^6\) is zero or negative at one of 201 radii between 0 and the widest ray. With the linear extension, the test stops at \(r_c\).

Chromatic Aberration#

The lateral chromatic aberration scales the ray of the red channel and of the blue channel. lens.chromatic_aberration_px gives the shift \(p_{ca}\) at the widest ray, in output pixels.

\[ c = \frac{p_{ca}}{f_x \sqrt{X_{max}^2 + Y_{max}^2}} \]
\[ (x, y)_{red} = (1 + c)(x, y), \qquad (x, y)_{green} = (x, y), \qquad (x, y)_{blue} = (1 - c)(x, y) \]

The model applies the scale only when the lens model is on.

Rolling Shutter#

Source: FAcresCameraModel::RowTransform and FAcresCameraModel::FromUnrealBody.

A rolling shutter reads the rows in sequence. The top row is first. The render is the image of the middle row. Row \(v\) has the time offset

\[ \Delta t_v = \left(v - \frac{H-1}{2}\right) \frac{T_{ro}}{H} \]

The camera rotates by \(\vec{\phi} = \vec{\omega}\,\Delta t_v\) in that time. With \(\theta = \lvert\vec{\phi}\rvert\), the unit axis \(\mathbf{a} = \vec{\phi}/\theta\) and its cross-product matrix \([\mathbf{a}]_\times\), the Rodrigues formula gives the rotation.

\[ R_v = I + \sin\theta\,[\mathbf{a}]_\times + (1 - \cos\theta)\,[\mathbf{a}]_\times^2 \]

The translation of the camera moves the ray by the translation divided by the assumed depth.

\[ \mathbf{t}_v = \frac{\mathbf{v}_c\,\Delta t_v}{Z_a} \]

With rolling_shutter.assumed_depth_m = 0 the translation is zero. The ray of the row in the frame of the render is

\[ \mathbf{d} = R_v \begin{pmatrix} x \\ y \\ 1 \end{pmatrix} + \mathbf{t}_v \]

Camera Motion#

The recorder reads the motion in the frame of the sample and holds it constant during the readout. \(\vec{\omega}\) is the angular velocity of the chassis. The velocity of the camera point includes the lever arm from the centre of mass:

\[ \mathbf{v}_{cam} = \mathbf{v}_{com} + \vec{\omega} \times (\mathbf{p}_{cam} - \mathbf{p}_{com}) \]

The recorder rotates the two vectors into the camera axes of Unreal: x forward, y right, z up. That frame is left-handed. The conversion to the optical frame is a reflection, thus the angular velocity changes its sign.

\[ \mathbf{v}_c = (v_y,\; -v_z,\; v_x), \qquad \vec{\omega}_{opt} = -(\omega_y,\; -\omega_z,\; \omega_x) \]

Each camera row of the session log contains the motion in the field rolling_shutter_motion_optical.

Projection and Sample#

Source: FAcresCameraModel::Process.

The ray \(\mathbf{d}\) projects into the render with the pinhole model. The model limits \(d_z\) to \(10^{-6}\) or more.

\[ u_r = f_r \frac{d_x}{d_z} + c_{xr}, \qquad v_r = f_r \frac{d_y}{d_z} + c_{yr} \]

The model limits \((u_r, v_r)\) to the render and takes a bilinear sample of one colour channel. The sample uses linear values. A table decodes each 8-bit value \(C = \text{byte}/255\) with the inverse sRGB curve:

\[ L = \begin{cases} C / 12.92 & C \le 0.04045 \\ \left(\dfrac{C + 0.055}{1.055}\right)^{2.4} & C > 0.04045 \end{cases} \]

Vignetting#

The key lens.vignetting.model selects the gain \(V\). The model applies it only when the lens model is on.

Model Gain
poly \(V = 1 + a_2 r_d^2 + a_4 r_d^4\), with the distorted radius \(r_d^2 = x_d^{*2} + y_d^{*2}\) of the output pixel
cos4 \(V = (1 + r^2)^{-2}\), with the undistorted radius. This is \(\cos^4\) of the ray angle.
none \(V = 1\)

The signal of the pixel is \(L' = L\,V\). The shipped values \(a_2 = -0.14\) and \(a_4 = 0.02\) give \(V = 0.88\) at \(r_d = 1\).

Image Sensor#

Source: FAcresCameraModel::Process.

Bayer Mosaic#

Each pixel measures one colour. The key photometric.cfa gives the colours of the top left 2 x 2 block. The colour of the pixel \((u, v)\) is the character at the index \((u \bmod 2) + 2 (v \bmod 2)\) of that key. The model samples the render along the ray of that colour, with its chromatic aberration scale.

Electrons and Noise#

The analogue gain is \(g = \max(I / I_{base},\, 0.001)\). At the gain \(g\), \(E_{fw}/g\) electrons give the full scale of the ADC.

\[ e = \max(L', 0)\,\frac{E_{fw}}{g}\,P(u, v) \]

\(P(u, v) = 1 + \sigma_{prnu}\,n\) is the pixel gain. The model draws it one time for each pixel at the start. The dark signal is \(e_{dark} = I_{dark}\,t_{exp}\). The model adds the shot noise, the dark noise and the read noise as one normal number. The variance of the shot noise is equal to its mean.

\[ \sigma = \sqrt{e + e_{dark} + \sigma_{read}^2} \]
\[ e' = \operatorname{clamp}\!\left(e + e_{dark} + \sigma\,n,\; 0,\; E_{fw}\right) \]

The output keeps the dark signal. No stage subtracts it.

ADC#

The ADC adds the black level \(B\) and scales the electrons. \(DN_{max} = 2^b - 1\).

\[ DN = B + e'\,\frac{(DN_{max} - B)\,g}{E_{fw}} \]
\[ Q = \operatorname{clamp}\!\left(\lfloor DN + 0.5 \rfloor,\; 0,\; DN_{max}\right) \]

The shipped values give \(E_{fw} / (DN_{max} - B) = 2.605\) electrons for each count at ISO 100.

With the noise off, the model does not compute electrons:

\[ DN = B + \operatorname{clamp}(L', 0, 1)\,(DN_{max} - B) \]

The ISO, the pixel gain and the dark signal then have no effect.

The next stage subtracts the black level and normalises the mosaic:

\[ M(u, v) = \frac{Q - B}{DN_{max} - B} \]

Random Numbers#

All random numbers come from PCG32 generators (FAcresPcg32) and the Box-Muller transform. Each row has its own generator, thus the result does not change with the number of threads.

Use Initial State Sequence
Pixel gain \(P\) seed + 3027 65536 + row
Noise of one frame seed + 3027 + (frame index + 1) x 1000003 row

In a row, an even column draws two uniform numbers \(u_1\) and \(u_2\) and uses \(n = \sqrt{-2 \ln u_1} \cos(2\pi u_2)\). The subsequent odd column uses \(n = \sqrt{-2 \ln u_1} \sin(2\pi u_2)\) of the same numbers. The frame index counts the samples that the recorder captured. A dropped sample moves the index of the subsequent frames.

Demosaic and Output#

The demosaic stage computes the two colours that a pixel did not measure. The key photometric.demosaic selects the filters. At the border the mosaic mirrors without a copy of the edge pixel.

The symbols are the values of the mosaic near the pixel. \(C\) is the pixel. \(S_h\) and \(S_v\) are the sums of the two horizontal and the two vertical neighbours. \(S_{2h}\) and \(S_{2v}\) are the sums of the two pixels at a distance of 2 in each direction. \(S_d\) is the sum of the four diagonal neighbours.

Malvar-He-Cutler Filters#

The default malvar uses the 5 x 5 filters of Malvar, He and Cutler (2004).

At a green pixel, \(X_h\) is the colour of the horizontal neighbours and \(X_v\) is the colour of the vertical neighbours.

\[ \begin{aligned} G &= C \\ X_h &= \tfrac{1}{8}\left(5C + 4 S_h - S_{2h} - S_d + \tfrac{1}{2} S_{2v}\right) \\ X_v &= \tfrac{1}{8}\left(5C + 4 S_v - S_{2v} - S_d + \tfrac{1}{2} S_{2h}\right) \end{aligned} \]

At a red pixel or a blue pixel, \(X_o\) is the opposite colour (blue at red, red at blue).

\[ \begin{aligned} G &= \tfrac{1}{8}\left(4C + 2 (S_h + S_v) - (S_{2h} + S_{2v})\right) \\ X_o &= \tfrac{1}{8}\left(6C + 2 S_d - \tfrac{3}{2} (S_{2h} + S_{2v})\right) \end{aligned} \]

Bilinear Filters#

With bilinear, a green pixel uses \(X_h = S_h / 2\) and \(X_v = S_v / 2\). A red pixel or a blue pixel uses \(G = (S_h + S_v) / 4\) and \(X_o = S_d / 4\).

Encode#

The model encodes each channel with the sRGB curve and a table of 16384 entries.

\[ C = \begin{cases} 12.92\,L & L \le 0.0031308 \\ 1.055\,L^{1/2.4} - 0.055 & L > 0.0031308 \end{cases} \]

The alpha channel is 255. The outputs are these.

Output Format
Session log camera/front_<time>.png, with the sample time in microseconds, and one row in camera.jsonl
Sensor stream One camera frame with the encoding bgr8. Refer to Sensor Stream.
HUD The image at half resolution

The image sensor model can be off while the lens model or the rolling shutter is on. The model then samples the three channels of each pixel and encodes them directly.

Timestamps and Latency#

The clock of the sensors is the physics step of 1/120 s. Refer to Time Stepping and Determinism.

Field of the Camera Row Meaning
physics_step, simulation_time_s The latest physics step when the recorder made the capture job
sample_time_s The time of that step since the start of the episode
available_time_s sample_time_s + delay_s (0.05 s)
delivery_time_s The time of the physics step that released the row. The PNG file exists at that time.
pose The pose of the capture component in the rendered frame
frame_index The index of the captured frame, which selects the noise
processing_ms The run time of the camera model for this image

The pose of the image is the pose of the rendered frame, not the pose of a physics step. In lockstep mode the simulation answers a step request only when the recorder has no sample in work. All images of the requested steps are then complete.

The modelled delay delay_s applies only to the rows of the session log. The sensor stream sends an image when the camera model completes it. The stamp of the frame is the simulation time of the sample.

The row time of the rolling shutter in episode.json is rolling_shutter_line_time_s \(= T_{ro}/H\). Row \(v\) has the exposure centre sample_time_s \(+ \Delta t_v\).

The Polaris Camera#

The Polaris has a Reolink RLC-1212A camera. The overlay sensors_polaris.json replaces the default values with the recorded calibration.

Intrinsics and Lens#

The values are the camera_info message that the vehicle published for the frame reolink_camera_optical: 896 x 512 px with a plumb_bob model.

Quantity Value
\(f_x, f_y\) 599.526 px, 592.378 px
\(c_x, c_y\) 440.130 px, 270.805 px
\(k_1, k_2, k_3\) -0.55389, 0.40842, -0.16673
\(p_1, p_2\) -0.0099785, 0.0044239
\(s_{min}\) and the limit radius 0.3, \(r_c = 0.954\)
Effective field of view 100.07° x 51.61°
Readout time \(T_{ro}\) 0.030 s

The recorded polynomial turns back at \(r = 0.96\), which is 57° from the axis and inside the corners of the image. The linear extension keeps the calibration inside the calibrated field and gives rays to the corners.

Mount#

The mount is relative to base_footprint, the middle of the rear axle on the ground. Calibration/Polaris/camera_extrinsic.py solved it from the LiDAR-camera calibration recording of 2 July 2026.

Parameter Value 1 σ (Bootstrap)
x (forward) 2.745 m 0.02 m
y (left) 0.606 m 0.40 m
z (up) 1.887 m 0.08 m
Roll -3.49° 0.5°
Pitch (down) 16.52° 0.4°
Yaw -2.63° 1.0°

The method uses 105 images of a marker target and the static LiDAR points on the target. The horizon in 12 recordings of the parked vehicle adds the roll reference. The distances of the target from the camera and from the LiDAR agree to 3 cm, with a scale of 1.002. The lateral position y is the weak parameter.

Comparison with the Real Camera#

Calibration/Polaris/camera_results.md compares rendered frames with 133 real frames at the same pose and time. The measure is the output of the SegFormer network that the vehicle uses.

Measure Value
Drivable IoU, pose correction from other frames of the run (median of the frames) 0.82
Drivable IoU at the refined poses, all pixels 0.79
Pixel accuracy, 9 classes, refined poses 0.75
Mean IoU, 9 classes, refined poses 0.37
Fréchet distance of the encoder features, real against simulated 208
Fréchet distance, real against real 4.2

The classes agree where the scene contains the objects. The encoder features show that the images are not the same as real images.

Ideal Camera Options#

These options remove parts of the model. The result is useful as ground truth.

Option Effect
-SensorCameraIdeal No lens model, no rolling shutter and no image sensor model. The output is the pinhole render at the output size.
-SensorCameraNoDistortion No lens model: no distortion, no principal point offset, no vignetting and no chromatic aberration.
-SensorCameraNoRollingShutter All rows use the time of the render.
-SensorCameraNoNoise No pixel gain variation and no noise. The Bayer mosaic, the ADC and the demosaic stay.
-SensorNoNoise No noise in all sensors. For the camera it is equal to -SensorCameraNoNoise.

The model is active when one of its three parts is on: the lens, the rolling shutter or the image sensor. The intrinsics in episode.json and in the sensor stream give zero distortion coefficients when the lens model is off.

Parameters#

The Camera Block of sensors.json#

The file is Acres/Content/Simulation/sensors.json. The option -SensorConfig= selects a different file.

Name Type Unit Default Description
seed integer 42 The seed of all sensor noise.
delay_s number s 0.05 The modelled delay of the rows in the session log.
camera.enabled boolean true Records the camera.
camera.hz number Hz 10 The sample rate, 1 to 120.
camera.width integer px 896 The width of the output, 64 to 4096.
camera.height integer px 512 The height of the output, 64 to 2160.
camera.hfov_deg number deg 90 The pinhole field of view that gives \(f_x\), 20 to 120.
camera.position_flu_m array of 3 m [0.6, 0, 1.65] The position in the mount frame, FLU.
camera.rpy_deg array of 3 deg [0, 0, 0] Roll, pitch and yaw of the mount.
camera.warmup_frames integer 2 \(N_w\), 0 to 16.
camera.temporal_history boolean true Keeps the temporal history of the capture.
camera.exposure.mode string volume volume, physical or copy.
camera.exposure.fstop number 8 The f-number \(N\).
camera.exposure.shutter_s number s 0.003 The exposure time \(t_{exp}\).
camera.exposure.iso number 100 The ISO \(I\).
Name Type Unit Default Description
lens.enabled boolean true The lens model.
lens.distortion_model string brown_conrady The only permitted value.
lens.k1, lens.k2, lens.k3 number -0.08, 0.012, 0 Radial coefficients, absolute value below 1.
lens.p1, lens.p2 number 0.0004, -0.0003 Tangential coefficients, absolute value below 0.05.
lens.principal_offset_px array of 2 px [3.5, -2.5] \((o_x, o_y)\), absolute value below 200.
lens.intrinsics object px none fx, fy, cx, cy of a calibrated camera.
lens.radial_min_slope number 0 \(s_{min}\), 0 to 1. The value 0 sets the linear extension off.
lens.vignetting.model string poly poly, cos4 or none.
lens.vignetting.a2, a4 number -0.14, 0.02 \(a_2, a_4\).
lens.chromatic_aberration_px number px 0.7 \(p_{ca}\), 0 to 5.
lens.render_scale number 1.0 \(s_r\), 0.5 to 2.
lens.render_margin number 0.02 \(m\), 0 to 0.2.
rolling_shutter.enabled boolean true The rolling shutter.
rolling_shutter.readout_s number s 0.025 \(T_{ro}\), 0 to 0.1.
rolling_shutter.assumed_depth_m number m 15 \(Z_a\), 0 to 1000.
photometric.enabled boolean true The image sensor model.
photometric.noise boolean true The pixel gain variation and the noise.
photometric.pixel_pitch_um number µm 2.8 The pixel pitch. The model only reports it.
photometric.full_well_e number e⁻ 10000 \(E_{fw}\), 500 to 10⁶.
photometric.read_noise_e number e⁻ 2.5 \(\sigma_{read}\), 0 to 100.
photometric.dark_current_e_per_s number e⁻/s 50 \(I_{dark}\), 0 to 10⁵.
photometric.prnu_std number 0.01 \(\sigma_{prnu}\), 0 to 0.1.
photometric.base_iso number 100 \(I_{base}\).
photometric.adc_bits integer bit 12 \(b\), 8 to 16.
photometric.black_level_dn number DN 256 \(B\), less than one quarter of the range.
photometric.cfa string RGGB RGGB, BGGR, GRBG or GBRG.
photometric.demosaic string malvar malvar or bilinear.

The names in the second table are relative to the camera block. The values are placeholders of a generic machine-vision camera. They are not measurements.

The Camera Block of sensors_polaris.json#

The Polaris loads sensors.json and then Acres/Content/Simulation/sensors_polaris.json. A key of the overlay replaces the same key of the base file. The option -PolarisSensorConfig= selects a different overlay. All other values stay those of sensors.json.

Name Type Unit Default Description
mount_frame string base_footprint The mount positions are relative to base_footprint.
camera.hz number Hz 10 The sample rate.
camera.width, camera.height integer px 896, 512 The size of the output.
camera.hfov_deg number deg 73.54 Not in use, because lens.intrinsics gives \(f_x\).
camera.position_flu_m array of 3 m [2.745, 0.606, 1.887] The solved mount position.
camera.rpy_deg array of 3 deg [-3.49, 16.52, -2.63] The solved mount rotation.
camera.warmup_frames integer 2 \(N_w\).
lens.intrinsics object px 599.526, 592.378, 440.130, 270.805 fx, fy, cx, cy.
lens.k1, lens.k2, lens.k3 number -0.55389, 0.40842, -0.16673 Radial coefficients.
lens.p1, lens.p2 number -0.0099785, 0.0044239 Tangential coefficients.
lens.radial_min_slope number 0.3 \(s_{min}\).
lens.chromatic_aberration_px number px 0.5 \(p_{ca}\).
lens.render_scale number 0.7 \(s_r\).
rolling_shutter.readout_s number s 0.03 \(T_{ro}\).

Command-Line Options#

The options replace the values of the files. Command-Line Options gives all options of the game.

Name Type Unit Default Description
-SensorCameraHz= number Hz file value The sample rate.
-SensorCameraWidth= integer px file value The width of the output.
-SensorCameraHeight= integer px file value The height of the output.
-SensorCameraHfov= number deg file value The pinhole field of view.
-SensorCameraWarmup= integer file value \(N_w\), limited to 0 to 16.
-SensorCameraNoHistory flag off Sets temporal_history to false.
-SensorCameraExposure= string file value volume, physical or copy.
-SensorCameraReadout= number s file value \(T_{ro}\).
-SensorCameraIdeal flag off Refer to Ideal Camera Options.
-SensorCameraNoDistortion flag off Sets the lens model off.
-SensorCameraNoRollingShutter flag off Sets the rolling shutter off.
-SensorCameraNoNoise flag off Sets the camera noise off.
-SensorCameraSelfTest flag off Writes the self-test of the model into the episode folder.
-SensorCameraMainFamily flag off Renders the capture in the view family of the main view. For diagnosis.
-SensorCameraShowFlags= string none Sets show flags of the capture, for example Bloom=0,Grain=0. For diagnosis.
-SensorNoCamera flag off Does not record the camera.

A camera setting that is different from the shipped files changes the label sensor_profile of the dataset to custom.

Self-Test#

With -SensorCameraSelfTest, FAcresCameraModel::RunSelfTest writes the folder camera_selftest into the episode folder.

File Contents
camera_selftest.json The intrinsics, the render geometry, the run time and 63 geometry samples with and without a test motion
raw_srgbNNN.u16 128 x 128 ADC counts of six flat grey levels, for noise statistics
checker_motion.png A checker pattern through the full model with the test motion

Code Map#

Item File Function
Configuration and limits Acres/Source/Acres/AcresCameraModel.cpp FAcresCameraModelConfig::Parse, ApplyCommandLine, Validate
Recorder configuration Acres/Source/Acres/AcresSensors.cpp FAcresSensorConfig::Load
Intrinsics, inverse table, render size, pixel gain Acres/Source/Acres/AcresCameraModel.cpp FAcresCameraModel::Initialize
Forward distortion and Jacobian Acres/Source/Acres/AcresCameraModel.cpp Distort
Rolling shutter Acres/Source/Acres/AcresCameraModel.cpp FAcresCameraModel::RowTransform, FromUnrealBody
Sample, vignetting, noise, ADC, demosaic Acres/Source/Acres/AcresCameraModel.cpp FAcresCameraModel::Process
Scene capture and exposure Acres/Source/Acres/AcresSensors.cpp FAcresSensorRecorder::Start
Scheduling and readback Acres/Source/Acres/AcresSensors.cpp FAcresSensorRecorder::TickCamera, EnqueueCameraCopy, CopyReadbackPixels
Outputs Acres/Source/Acres/AcresSensors.cpp FAcresSensorRecorder::FinishCapture
Random numbers Acres/Source/Acres/AcresSensors.cpp FAcresPcg32
Sensor profile Acres/Source/Acres/AcresRenderTier.cpp FAcresRenderTier::PinSensors
Polaris lens for the calibration tools Calibration/Polaris/reolink_lens.py
Polaris mount solution Calibration/Polaris/camera_extrinsic.py
Rendered frames at real poses Calibration/Polaris/camera_render.py, camera_compare.py, camera_evaluate.py

Limitations#

  • The signal is the tone-mapped render. The noise does not follow the scene radiance or the exposure, except for the ISO gain and the dark signal.
  • The rolling shutter uses one render, one depth \(Z_a\) and a constant motion. It does not show the parallax of near objects, moving objects or new visible areas.
  • The exposure time adds no motion blur.
  • The camera has no white balance, no colour correction matrix, no noise reduction and no sharpening.
  • The pose of an image is the pose of the rendered frame. Without lockstep mode it is not the pose of a fixed physics step.
  • The default camera uses placeholder values. Only the Polaris camera has a calibration.
  • The lateral position of the Polaris camera has an uncertainty of 0.40 m (1 σ).
  • The game has one camera for each vehicle.

References#

  • Brown, D. C. (1966). Decentering distortion of lenses. Photogrammetric Engineering, 32(3), 444-462.
  • Conrady, A. E. (1919). Decentred lens-systems. Monthly Notices of the Royal Astronomical Society, 79(5), 384-390.
  • Bradski, G. (2000). The OpenCV library. Dr. Dobb's Journal of Software Tools, 25(11), 120-125.
  • Malvar, H. S., He, L.-W., and Cutler, R. (2004). High-quality linear interpolation for demosaicing of Bayer-patterned color images. Proceedings of IEEE ICASSP 2004, vol. 3, 485-488.
  • IEC 61966-2-1 (1999). Multimedia systems and equipment - Colour measurement and management - Part 2-1: Default RGB colour space - sRGB. International Electrotechnical Commission.
  • O'Neill, M. E. (2014). PCG: A family of simple fast space-efficient statistically good algorithms for random number generation. Technical report HMC-CS-2014-0905, Harvey Mudd College.
  • Box, G. E. P., and Muller, M. E. (1958). A note on the generation of random normal deviates. The Annals of Mathematical Statistics, 29(2), 610-611.