<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.8.5">Jekyll</generator><link href="https://noorvir.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://noorvir.github.io/" rel="alternate" type="text/html" /><updated>2019-08-27T22:50:55+00:00</updated><id>https://noorvir.github.io/feed.xml</id><title type="html">Noorvir Aulakh</title><subtitle></subtitle><author><name>Noorvir Aulakh</name></author><entry><title type="html">Future Of Robotics</title><link href="https://noorvir.github.io/Future-of-Robotics.html" rel="alternate" type="text/html" title="Future Of Robotics" /><published>2019-07-01T00:00:00+00:00</published><updated>2019-07-01T00:00:00+00:00</updated><id>https://noorvir.github.io/Future-of-Robotics</id><content type="html" xml:base="https://noorvir.github.io/Future-of-Robotics.html">&lt;p&gt;The Role of Robotics and Automation Systems in Shaping the Future of the Manufacturing Industry&lt;/p&gt;

&lt;p&gt;The advances in computing and consumer electronics over the last twenty years along with the emergence of the internet have catalysed a new wave in technological innovation and enterprise¬¬¬. The commoditisation of electronics and computing hardware has engendered a rapid shift in what was once considered high-end, industrial grade technology to the public domain. It has allowed new companies to spring up and innovate at an unprecedented rate in what could perhaps be called the 21st century start-up model. 
The combination of these factors has resulted in the emergence of the Internet of Things (IoT) which is the principal underlier for what is being termed as Industry 4.0, or the 4th Industrial revolution . Industry 4.0 is the view of the modern-day factory as network of smart, interconnected nodes, wherein the nodes take the form of any equipment that adds value to the final product and that can be made accessible through the network (the IoT).&lt;/p&gt;

&lt;p&gt;The overarching goal at a factory is to extract as much value as possible from as few resources as possible. To this end, production lines are primarily optimized to minimize two variables: takt-time  and down-time . While there is a whole range of operational and workplace management techniques (including 5S and six sigma) aimed at addressing the former, the latter is much more prone to vicissitudes of the supply chain and therefore much harder to optimize. 
This issue is finally becoming addressable through Industry 4.0; real-time tracking information on incoming shipments and their estimated arrival times, allows the production line to schedule operation accordingly. A similar principle can be applied to part movement within a factory – network connected supply carts communicate with the production line to relay information about required parts and their estimated delivery time to the associated work-stations line-side. The tracking of such supply-carts also allows for real-time optimization of the flow of traffic within a factory . While this system is only at its nascent stages of implementation, as the infrastructure proliferates, the paradigm for part movement and warehousing itself will undergo an upheaval - replacing human driven carts and fork-lifts in favour of their autonomous counterparts.&lt;/p&gt;

&lt;p&gt;This trend was first evident in the technology developed by Kiva Systems for automated control of warehousing, and its utility made equally clear when Amazon Inc. bought the company for $$775 million, in March 2012  . While Amazon’s decision to dedicate the Kiva technology for use solely in Amazon warehouses could have been a step back for the field, the 21st century start-up model jumped in to bridge the gap; companies like ClearPath Robotics, whose early core technology was based on consumer grade electronics, are working on systems that automate part movement across factory floors and warehouses, with the aim of making Amazon’s proprietary technology an industry standard  . 
The economic benefit of automating these well-defined tasks in controlled environments (warehouses and factory floors) makes it easy to extrapolate the adoption of this technology over the next ten years to a stage where it would be unfeasible for a mass-producing manufacturer not utilise it. What is much harder to address and predict is the automation of tasks that are equally mundane which nevertheless require much greater dexterity and skilled labour to accomplish. Installing parts onto the chassis of a car for instance requires a level of cognition that is currently hard to attain in robots. These tasks access a specific facet of human learning termed as reinforcement learning.&lt;/p&gt;

&lt;p&gt;Reinforcement learning (RL) is so innate to humans that we are scarcely aware of it. We subconsciously learn skills such as walking or catching a ball which are incredibly complex to encode in a robot. Empowered by the advances in computing capability over the recent years, however, we are beginning to apply RL to machines and producing results which surprise even their creators. Google Deep Mind’s AlphaGo computer (based on RL) for instance, beat the South Korean champion Lee Sedol at the Chinese game of GO which is known to be unsolvable using traditional brute force methods (based on current day computing capabilities)  . This approach might prove pivotal in the next generation of automation technologies in manufacturing.&lt;/p&gt;

&lt;p&gt;Research on Deep Reinforcement Learning applied to robot manipulation utilises similar techniques as Google’s AlphaGo but applies them to allow robots to understand the dynamics of the physical world and to interact with it in a dextrous manner  . Such techniques would allow robots to accomplish our earlier example of installing parts on a car, which might have been deemed undoable.&lt;/p&gt;

&lt;p&gt;Over the next ten years, it is improbable that robots will be capable to human level dexterity and understanding of the physical world. What is more likely is the emergence of collaborative human-robot environments where machines in a sense act as a human’s assistants or counterpart in performing repetitive tasks. Baxter by Rethink Robotics is one such robot which is at the leading edge of this trend .  Although Baxter is currently based on learning through human demonstration and does not as such utilise RL, with its focus on human-friendly collaborative interaction and two high degree of freedom dextrous arms, it is ideally poised to be ubiquitous on factory floors over the coming decade and capitalise on the advancements in RL .&lt;/p&gt;

&lt;p&gt;Lastly, an aspect of robotics that is often overlooked, since it takes a step away from the physical world, yet, is equally important in manufacturing, is machine learning applied to predictive analysis and optimisation. Machine learning algorithms for instance could make time consuming diagnostics on faulty components unnecessary by suggesting a “most likely” failure mode based on some learned parameters. In the remanufacturing department of a car company, this could mean not having to physically take apart faulty components to identify the underlying problem, thereby saving material and labour costs.&lt;/p&gt;

&lt;p&gt;Machine learning could also be applied to optimise the performance of the factory and reorganise workflow to save resources. This application was demonstrated by Google when they used their DeepMind AI to reduce the power consumption of their data-centres by 40% . Such demonstrable economic benefit is a sure-shot indicator of the wide-spread proliferation of this technology as the new generation of engineers and scientists move into industry over the next ten years.&lt;/p&gt;

&lt;p&gt;Any discussion of automation is incomplete and irresponsible unless it addresses the potential impact on the human work-force. Since the first industrial revolution, advancements in technology that questioned the position of a human work-force have engendered controversy and in most cases, been dismissed as the work of luddites, and to some degree, rightly so. With Industry 4.0 and the advancements in AI, the situation this time is different. While the short-term (10-year timescale) role of robotics and automation is expected to be hugely economically beneficial, in the long term, it could lead to a range of issues, from inequality to unemployment and an eventual economic downturn . In creating a ‘Partnership on AI’ leading tech-corporations such as Facebook, Google and Amazon, realise this eventuality and are conscious about ameliorating it. Given that the pace of legislation is rarely able to keep up with industrial advancement (especially when vast economic benefits are to be reaped), partnerships such are these are our best bet to prevent a reversion to the social divide and stagnant economy of the feudal ages.&lt;/p&gt;</content><author><name>Noorvir Aulakh</name></author><summary type="html">The Role of Robotics and Automation Systems in Shaping the Future of the Manufacturing Industry</summary></entry><entry><title type="html">TSDF-Fusion with a Robot Arm</title><link href="https://noorvir.github.io/TSDF-Fusion-with-a-Robot-Arm.html" rel="alternate" type="text/html" title="TSDF-Fusion with a  Robot Arm" /><published>2018-03-31T00:00:00+00:00</published><updated>2018-03-31T00:00:00+00:00</updated><id>https://noorvir.github.io/TSDF-Fusion-with-a%20-Robot-Arm</id><content type="html" xml:base="https://noorvir.github.io/TSDF-Fusion-with-a-Robot-Arm.html">&lt;p&gt;&lt;a href=&quot;http://people.inf.ethz.ch/moswald/publications/&quot;&gt;&lt;img src=&quot;assets/img/tsdf/3D_recon.png&quot; alt=&quot;alt&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;The remarkable progress in 2D computer vision and its innumerable applications in recent years has meant that there is a trove of accessible reference material for newcomers to the field. Unfortunately, this is not the case for the some of the more classical (sans deep-deep-neural-nets) approaches in 3D computer vision. These techniques nevertheless remain vitally important to a variety of fields and, in my opnion, will make a come-back in hybrid form with neural-nets in the near future.&lt;/p&gt;

&lt;p&gt;In this blog post we’ll talk about one of these fundamental techniques in 3D computer vision: 3D-reconstruction using Truncated Signed Distance Function (TSDF) Fusion. Concretely, we’ll cover the practical details of combining depth images gathered from multiple known camera positions into a 3D surface reconstruction. We’ll also address some of the technical challenges you might come across when dealing with depth images, and present a slightly unconventional way of dealing with them (in 2D). Let’s break down this process into the following components and tackle them one by one:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;#why-tsdf-?&quot;&gt;Why TSDF?&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#physical-setup&quot;&gt;Physical Setup&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#camera-calibration&quot;&gt;Camera Calibration&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#fusion&quot;&gt;Fusion&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#experiments-with-depth-images&quot;&gt;Experiments with Depth Images&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;#applications&quot;&gt;Applications&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each section can be read independently without the need to read any of the others as long as you have a general understanding of the concepts therein.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;N.B.&lt;/em&gt;&lt;/strong&gt; This is not a blog post about implementing TSDF-Fusion. Instead, we’ll focus on practical considerations and setup. You can find excellent open-source &lt;a href=&quot;https://github.com/ethz-asl/voxblox&quot;&gt;CPU&lt;/a&gt; and &lt;a href=&quot;https://github.com/andyzeng/tsdf-fusion&quot;&gt;GPU&lt;/a&gt; implementations online.&lt;/p&gt;

&lt;h2 id=&quot;why-tsdf&quot;&gt;Why TSDF?&lt;/h2&gt;

&lt;p&gt;It is first and foremost important to ask yourself whether TSDF-Fusion is what you need - which implies understanding what it actually is. Technically, in computer-graphics, &lt;a href=&quot;https://en.wikipedia.org/wiki/Signed_distance_function&quot;&gt;SDF&lt;/a&gt; refers to a representation of a 3D surface on a volumetric (&lt;a href=&quot;https://en.wikipedia.org/wiki/Voxel&quot;&gt;voxel&lt;/a&gt;) grid, where the value of the function at each voxel approximates its distance from a 3D surface.&lt;/p&gt;

&lt;p&gt;TSDF-Fusion can be used as a component in a  &lt;a href=&quot;https://en.wikipedia.org/wiki/Simultaneous_localization_and_mapping&quot;&gt;Simultaneous Localisation and Mapping&lt;/a&gt; (SLAM) algorithm, which is in-fact the case (amongst other) in the original &lt;a href=&quot;https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/ismar2011.pdf&quot;&gt;KinectFusion paper&lt;/a&gt;. In this article however, we ignore the “Localisation” part and focus solely on the fusion of depth images (where each pixel represents the distance to a 3D point) captured from a camera with known extrinsics (6 DoF pose relative to some reference coordinate system).&lt;/p&gt;

&lt;h2 id=&quot;physical-setup&quot;&gt;Physical Setup&lt;/h2&gt;

&lt;p&gt;In our case, the set-up consists of a depth camera mounted on a robot-arm. This allows us to determine the camera pose through forward kinematics.&lt;/p&gt;

&lt;div&gt;

&lt;table class=&quot;table_align&quot; style=&quot;width:100%&quot;&gt;
    &lt;td class=&quot;table_align&quot;&gt;
        &lt;img src=&quot;assets/img/tsdf/physical_setup_small.png&quot; alt=&quot;Physical Setup&quot; /&gt;
    &lt;/td&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 1: Physical Setup. The pointer is used to localise the position of the calibration rig as discussed in the next section.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We use a PMD &lt;a href=&quot;https://pmdtec.com/picofamily/flexx/&quot;&gt;pico flexx&lt;/a&gt; depth camera fixed onto the robot-arm using a 3D-printed mount. The pico flexx uses time-of-flight technology, which although noisy from one frame to the next, yields relatively accurate point-clouds. The table below compares the pico flexx against other depth cameras commonly used in the research community.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Camera Name&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Technology&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Depth Range (m)&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Frame Rate (fps)&lt;/th&gt;
      &lt;th style=&quot;text-align: center&quot;&gt;Size (mm)&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Pico Flexx&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Time-of-Flight&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;0.1 - 7&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;5 - 45&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;68 x 17 x 7.35&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Intel Realsense D435&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;Active IR&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;0.2 - 10&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;30 - 90&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;90 x 25 x 25&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Primesense  Carmine&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;-&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;0.3 - 3.5&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;30&lt;/td&gt;
      &lt;td style=&quot;text-align: center&quot;&gt;180 x 25 x 35&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Our use case in particular requires operating between 0.15 - 3 meters from the object surface, which the pico flexx is ideal for. It however does not come with an in-built RGB camera, so you’d need to manually calibrate a separate RGB camera (which might be a deal breaker for some).&lt;/p&gt;

&lt;h2 id=&quot;camera-calibration&quot;&gt;Camera Calibration&lt;/h2&gt;

&lt;p&gt;Whether or not you’re using the pico flexx, it is essential to have a good intrinsics and extrinsics calibration of your camera. This amounts to estimating the &lt;script type=&quot;math/tex&quot;&gt;\mathbf{K}&lt;/script&gt; and &lt;script type=&quot;math/tex&quot;&gt;\mathbf{T}&lt;/script&gt; matrix in the camera projection equation below, which transforms a 3D point &lt;script type=&quot;math/tex&quot;&gt;\mathbf{w}&lt;/script&gt; (specified in homogenous coordinates) from some arbitrary reference frame to 2D pixel coordinates &lt;script type=&quot;math/tex&quot;&gt;\mathbf{u}&lt;/script&gt; in the camera frame.&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;% &lt;![CDATA[
\lambda

\begin{bmatrix}

u \\
v \\
1

\end{bmatrix}

=

\underbrace{

     \begin{bmatrix}
     \phi_x &amp; \gamma &amp; \delta_x  &amp; 0  \\
     0 &amp; \phi_y &amp; \delta_y  &amp; 0  \\
     0 &amp; 0 &amp; 1  &amp; 0  \\

     \end{bmatrix}
}_\mathbf{K}

\underbrace{

\begin{bmatrix}
\omega_{11} &amp; \omega_{12} &amp; \omega_{13}  &amp; \tau_x  \\
\omega_{21} &amp; \omega_{22} &amp; \omega_{23}  &amp; \tau_y  \\
\omega_{31} &amp; \omega_{32} &amp; \omega_{33}  &amp; \tau_z  \\
0 &amp; 0 &amp; 0  &amp; 1  \\

\end{bmatrix}
}_\mathbf{T}

\begin{bmatrix}

x \\
y \\
z \\
1

\end{bmatrix}
\tag{1}\label{eq:one} %]]&gt;&lt;/script&gt;

&lt;h3 id=&quot;intrinsics&quot;&gt;Intrinsics&lt;/h3&gt;

&lt;p&gt;Intrinsics calibration refers to the estimation of the &lt;a href=&quot;http://www.vision.caltech.edu/bouguetj/calib_doc/htmls/parameters.html&quot;&gt;camera matrix&lt;/a&gt; &lt;script type=&quot;math/tex&quot;&gt;\mathbf{K}&lt;/script&gt; which accounts for the projection model of the pinhole-camera. It takes into account the the distance and alignment of the image plane relative to the camera optical center (&lt;script type=&quot;math/tex&quot;&gt;F_c&lt;/script&gt; in Figure 2). A 3D point &lt;script type=&quot;math/tex&quot;&gt;\mathbf{w}=(x, y, z)&lt;/script&gt; in camera coordinate frame &lt;script type=&quot;math/tex&quot;&gt;F_c&lt;/script&gt;,  projects onto the camera image plane at pixels &lt;script type=&quot;math/tex&quot;&gt;\mathbf{u}=(u, v)&lt;/script&gt;.&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;u = \frac{\phi_x x + \gamma y}{z} + \delta_x

\tag{2} \label{eq:two}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;v = \frac{\phi_y y}{z} + \delta_y

\tag{3}\label{eq:three}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Applying the camera matrix in equation 1 above normalises the camera model: yields focal length (&lt;script type=&quot;math/tex&quot;&gt;\phi_x&lt;/script&gt; and &lt;script type=&quot;math/tex&quot;&gt;\phi_y&lt;/script&gt;) 1, offsets (&lt;script type=&quot;math/tex&quot;&gt;\delta_x&lt;/script&gt; and &lt;script type=&quot;math/tex&quot;&gt;\delta_y&lt;/script&gt;) the location of the optical centre relative to the top left corner of the image, and corrects for manufacturing “defects” such as irregular photoreceptor spacing and skew (&lt;script type=&quot;math/tex&quot;&gt;\gamma&lt;/script&gt;).&lt;/p&gt;

&lt;div&gt;

&lt;table class=&quot;table_align&quot; style=&quot;width:100%&quot;&gt;
    &lt;td class=&quot;table_align&quot;&gt;
        &lt;img src=&quot;assets/img/tsdf/pin_hole_model.png&quot; alt=&quot;Pinhole camera model&quot; /&gt;
    &lt;/td&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 2: Pinhole Camera Model.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In addition to the projection parameters, the formation of an image on the camera’s pixel array is also affected by the “focusing” effect of the lens - which a pinhole model ignores. These can be summarised with a set of radial &lt;script type=&quot;math/tex&quot;&gt;(k_1, k_2, k_3)&lt;/script&gt; and tangential &lt;script type=&quot;math/tex&quot;&gt;(p_1, p_2)&lt;/script&gt; &lt;a href=&quot;https://www.mathworks.com/help/vision/ug/camera-calibration.html#bu0nj3f&quot;&gt;distortion coefficients&lt;/a&gt;. Distortion correction is applied after normalising (perspective projection &lt;script type=&quot;math/tex&quot;&gt;\mathbf{x \rightarrow}\mathbf{x'}&lt;/script&gt;) the 3D point but before the pin-hole projection (the action of &lt;script type=&quot;math/tex&quot;&gt;\mathbf{K}&lt;/script&gt;).&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\begin{bmatrix}

x' \\
y' \\

\end{bmatrix}
=
\begin{bmatrix}

x/z \\
y/z \\

\end{bmatrix}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\hat x = \underbrace{
          (1 + k_1 r^2 + k_2 r^4 + k_3 r^6) x'}_{Radial Component} +
          \underbrace{
          2 p_1 x' y' + p_2 (r^2+ 2 x'^2)}_{Tangential Component}

\tag{4}\label{eq:four}&lt;/script&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\hat y = \overbrace{
          (1 + k_1 r^2 + k_2 r^4 + k_3 r^6) y'} +
          \overbrace{
          p_1 (r^2+ 2 x'^2) + 2 p_2 x' y'}

\tag{5}\label{eq:five}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Where &lt;script type=&quot;math/tex&quot;&gt;r^2 = x^2 + y^2&lt;/script&gt;. You can now get the pixel projections &lt;script type=&quot;math/tex&quot;&gt;(u, v)&lt;/script&gt; by replacing &lt;script type=&quot;math/tex&quot;&gt;(x, y)&lt;/script&gt; in equation &lt;script type=&quot;math/tex&quot;&gt;\eqref{eq:two}&lt;/script&gt; and &lt;script type=&quot;math/tex&quot;&gt;\eqref{eq:three}&lt;/script&gt; by &lt;script type=&quot;math/tex&quot;&gt;(\hat x, \hat y)&lt;/script&gt;. The problem of intrinsics calibration therefore is to estimate the following parameters: {&lt;script type=&quot;math/tex&quot;&gt;\phi_x, \phi_y, \delta_x, \delta_y, \gamma, k_1, k_2, k_3, p_1, p_2&lt;/script&gt;}.&lt;/p&gt;

&lt;p&gt;Intrinsics calibration is a fairly standard procedure; you can find a step-by-step tutorial &lt;a href=&quot;https://docs.opencv.org/3.4.3/dc/dbb/tutorial_py_calibration.html&quot;&gt;here&lt;/a&gt;.
Most calibration algorithms presume images taken from multiple, varied vantage points where known world coordinates can be identified reliably (such as a the corners of chessboard pattern). On the pico flexx, we can do this by accessing the intensity image &lt;em&gt;Figure 3b&lt;/em&gt;. The pixel values here represent the magnitude of “excitement” of the photoreceptors, which you’ll need to normalise into a range that libraries such as OpenCV expect (&lt;code class=&quot;highlighter-rouge&quot;&gt;CV_8U&lt;/code&gt;, &lt;code class=&quot;highlighter-rouge&quot;&gt;CV_16U&lt;/code&gt;, etc.)&lt;/p&gt;

&lt;div&gt;

&lt;table class=&quot;table_align&quot;&gt;
    &lt;tr&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/depth_image.png&quot; alt=&quot;Depth Image&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(a)&lt;/i&gt;
        &lt;/td&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/calib_pattern_intensity.png&quot; alt=&quot;Intensity Image&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(b)&lt;/i&gt;
        &lt;/td&gt;
    &lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 3: (a) A depth image taken from the pico flexx, (b) An intensity image of a calibration target (chessboard) with the detected corners highlighted and sorted.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Once you have a set of calibration &lt;code class=&quot;highlighter-rouge&quot;&gt;images&lt;/code&gt;, the procedure can be summarised in the following pseudo-code snippet.&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;c&quot;&gt;# Intrinsics Calibration&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;points_3d&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;zeros&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;([&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;board_size&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;**&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;3&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;points_3d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[:,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;mgrid&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(:&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;board_size&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;board_size&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;T&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;reshape&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;square_size&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;image&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;images&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;points_2d&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cv2&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;findChessboardCorners&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;image&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;board_size&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;board_size&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;points_2d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;img_points&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;world_points&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;points_3d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;cam_matrix&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;dist_coeffs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;rvecs&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;tvecs&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cv2&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;calibrateCamera&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;world_points&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;img_points&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;c&quot;&gt;# Optional Extra Extrinsics Calibration&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;T_cb_cams&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[]&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;world_point&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;img_point&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;zip&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;world_points&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;img_points&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;rvec_cb_cam&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;tvec_cb_cam&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cv2&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;solvePnPRansac&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;world_point&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;img_point&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

    &lt;span class=&quot;c&quot;&gt;# Convert to 4x4 transformation matrix&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;T&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;combine&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cv2&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Rodrigues&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;tvec_cb_cam&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;T_cb_cams&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;append&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;T&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;extrinsics&quot;&gt;Extrinsics&lt;/h3&gt;

&lt;p&gt;Unlike the case for a SLAM problem, we do not want to estimate the global pose of the camera at every time step. Instead, we want to estimate the extrinsic transformation of the camera’s optical center relative to its mounting point on the robot arm.&lt;/p&gt;

&lt;p&gt;At first glance, having a 3D CAD model of the camera-mount might appear to obviate the need for such a calibration. Unfortunately, in practice even small misalignments compound as a function of distance when rotations are involved, as visualised by &lt;em&gt;Figure 4&lt;/em&gt; below.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/rotation_compounding.png&quot; alt=&quot;alt The concept behind TSDF fusion &quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 4: The compounding effect of small misalignments on rays of light.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The &lt;code class=&quot;highlighter-rouge&quot;&gt;cv2.calibrateCamera&lt;/code&gt; method from the code snippet above simultaneously solves for both intrinsics and extrinsics, which are returned in axis-angle representation (&lt;code class=&quot;highlighter-rouge&quot;&gt;rvecs&lt;/code&gt;, &lt;code class=&quot;highlighter-rouge&quot;&gt;tvecs&lt;/code&gt;). In practice you get slightly better results if you carry out a second calibration of the extrinsics. In particular, you can now use &lt;a href=&quot;https://en.wikipedia.org/wiki/Random_sample_consensus&quot;&gt;RANSAC&lt;/a&gt; and ignore the possible effect of outliers.&lt;/p&gt;

&lt;h3 id=&quot;extrinsics-kinematic-chain&quot;&gt;Extrinsics Kinematic Chain&lt;/h3&gt;

&lt;p&gt;Solving the extrinsics optimisation gives us the 6 DoF transformation between the camera and the origin of the chessboard pattern &lt;script type=&quot;math/tex&quot;&gt;T^{cb}_{cam}&lt;/script&gt;. We still do not know how the camera (which is so far floating around in space) is orientated with respect to the robot arm &lt;script type=&quot;math/tex&quot;&gt;T^{cam}_{ra}&lt;/script&gt;. The missing link is the pose of the chessboard with respect to the global (robot) reference frame &lt;script type=&quot;math/tex&quot;&gt;T^{cb}_W&lt;/script&gt;. Once we know this transformation, estimating the camera pose with respect to this reference frame simply amounts to a chain of rigid transformations:&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;T^{cam}_{W} = \underbrace{
               T^{ra}_W \cdot T^{cb}_{ra}
               }_{T^{cb}_{W }}
               \cdot T^{cam}_{cb}&lt;/script&gt;

&lt;p&gt;Where the notation &lt;script type=&quot;math/tex&quot;&gt;T^a_b \cdot \vec{v}&lt;/script&gt; denotes the rigid 6-DoF transformation required to transform a vector &lt;script type=&quot;math/tex&quot;&gt;\vec{v}&lt;/script&gt; from reference frame &lt;script type=&quot;math/tex&quot;&gt;a&lt;/script&gt; to an arbitrary frame &lt;script type=&quot;math/tex&quot;&gt;b&lt;/script&gt;.&lt;/p&gt;

&lt;p&gt;You could either have a (very) precise measurement of &lt;script type=&quot;math/tex&quot;&gt;T^{cb}_W&lt;/script&gt; by placing the pattern at a pre-calibrated position, or measure &lt;script type=&quot;math/tex&quot;&gt;T^{cb}_{ra}&lt;/script&gt; using the robot as a pointing device. In this case, we mount a 3D printed pointer on the robot and point it at the origin of the chessboard pattern, while making sure to align the en-effector to match the coordinate axes of the chessboard.&lt;/p&gt;

&lt;div&gt;
&lt;table class=&quot;table_align&quot; style=&quot;width:100%&quot;&gt;
&lt;td class=&quot;table_align&quot;&gt;
    &lt;img src=&quot;assets/img/tsdf/calib_setup.png&quot; alt=&quot;Calibration Setup&quot; /&gt;
&lt;/td&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 5: Calibration setup. The coordinates of the chessboard &lt;script type=&quot;math/tex&quot;&gt;T^{cb}_W&lt;/script&gt; can be determined using a pointer with known forward kinematics &lt;script type=&quot;math/tex&quot;&gt;T_{ra}&lt;/script&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Practical Considerations:&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Make sure to image the chessboard from a range of view points.&lt;/li&gt;
  &lt;li&gt;If you have a lot of lens distortion, you can either use the fisheye calibration method in OpenCV, or you could only consider a central crop of the depth image. This is the case for the pico flexx, the central crop works well but the edges show large distortions.&lt;/li&gt;
  &lt;li&gt;In case you crop your images, remember to compensate the offset parameters &lt;script type=&quot;math/tex&quot;&gt;(\delta_x, \delta_y)&lt;/script&gt; in the Camera Matrix.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We now finally have all the ingredients to start fusing some TSDFs!&lt;/p&gt;

&lt;h2 id=&quot;fusion&quot;&gt;Fusion&lt;/h2&gt;

&lt;p&gt;While we’re not going to go into the implementation details of TSDF-Fusion, let’s have a quick theoretical overview to understand what’s happening under-the-hood, so that you can debug when things go awry.&lt;/p&gt;

&lt;p&gt;As we briefly discussed in the first section, our aim is to combine a number of depth images taken from known camera poses into a 3D reconstruction.  Since images taken from different vantage points might not align exactly (&lt;em&gt;Figure 6b&lt;/em&gt;), we take the weighted average of multiple independent surface measurements. The &lt;a href=&quot;https://graphics.stanford.edu/papers/volrange/volrange.pdf&quot;&gt;original paper&lt;/a&gt; introducing the technique gives a relatively comprehensible explanation.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/tsdf_concept.png&quot; alt=&quot;alt The concept behind TSDF fusion &quot; /&gt;
&lt;em&gt;Figure 6: Visualing the concept behind averaging two measured surfaces to fuse them into one. Curless, B., &amp;amp; Levoy, M. (1996)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The &lt;script type=&quot;math/tex&quot;&gt;T&lt;/script&gt; in TSDF becomes relevant when computing such a surface representation from a collection of “noisy” depth images. It has to do with the “truncation distance” &lt;script type=&quot;math/tex&quot;&gt;\epsilon&lt;/script&gt; behind a surface, within which we assume a depth measurement to originate from it. It is introduced to minimise the possibility of surfaces interacting when imaged from opposite directions, which becomes particularly important for relatively thin objects. &lt;em&gt;Figure 7&lt;/em&gt; visualises this concept, where the truncation margin around the inner side of the surface is highlighted in red.&lt;/p&gt;

&lt;!-- | ![alt The concept behind TSDF fusion ][tsdf_truncation_concept] |
|:--:| --&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/tsdf_truncation_concept.png&quot; alt=&quot;alt Truncation concept &quot; /&gt;
&lt;em&gt;Figure 7: Each voxel &lt;script type=&quot;math/tex&quot;&gt;v&lt;/script&gt; in a discrete voxel-grid contains its distance &lt;script type=&quot;math/tex&quot;&gt;d^v&lt;/script&gt; from the nearest surface. Depth measurement from the camera a truncated to a distance &lt;script type=&quot;math/tex&quot;&gt;\epsilon&lt;/script&gt; from the surface.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Each voxel &lt;script type=&quot;math/tex&quot;&gt;v&lt;/script&gt; on the grid contains its distance &lt;script type=&quot;math/tex&quot;&gt;d^v&lt;/script&gt; from the nearest surface. One way to compute the &lt;script type=&quot;math/tex&quot;&gt;d^v&lt;/script&gt; is to iterate through the entire voxel-grid, check whether each voxel is visible in a given camera frame and calculate:&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;d^v = I_d - v_{cam}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;where &lt;script type=&quot;math/tex&quot;&gt;v_{cam}&lt;/script&gt; is the distance of the voxel from the camera.&lt;/p&gt;

&lt;p&gt;With that in mind, the essence of the TSDF algorithm is summarised in the following four equations:&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\renewcommand{\vec}[1]{\mathbf{#1}}

D(\vec{x}) = \frac{\sum w_{i}(\vec{x}) d_{i}(\vec{x})}
			{\sum w_{i}(\vec{x})}

\tag{6}\label{eq:six}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\renewcommand{\vec}[1]{\mathbf{#1}}

W(\vec{x}) =  w_i(\vec{x})

\tag{7}\label{eq:seven}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\renewcommand{\vec}[1]{\mathbf{#1}}

D_{i+1}(\vec{x}) = \frac{W_i(\vec{x}) D_i(\vec{x}) + w_{i+1}(\vec{x})d_{i+1}(\vec{x})}{W_i(\vec{x}) + w_{i+1}(\vec{x})}

\tag{8}\label{eq:eight}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;\renewcommand{\vec}[1]{\mathbf{#1}}

W_{i+1}(\vec{x}) = W_{i}(\vec{x}) + w_{i+1}(\vec{x})

\tag{9}\label{eq:nine}&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Equation &lt;script type=&quot;math/tex&quot;&gt;\eqref{eq:six}&lt;/script&gt; gives the combination rule for the cumulative signed-distance function &lt;script type=&quot;math/tex&quot;&gt;D(\mathbf{x})&lt;/script&gt; that we’re trying to estimate where:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;script type=&quot;math/tex&quot;&gt;d_i(\mathbf{x})&lt;/script&gt; is the signed-distance map at each time-step (computed from the depth image).&lt;/li&gt;
  &lt;li&gt;&lt;script type=&quot;math/tex&quot;&gt;w_i(\mathbf{x})&lt;/script&gt; is the weight function for the current time-step.&lt;/li&gt;
  &lt;li&gt;&lt;script type=&quot;math/tex&quot;&gt;W(\mathbf{x})&lt;/script&gt; is the cumulative weight function - it takes into account all the measurements so far.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The weight function serves two roles:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;It weighs successive depth measurements against the cumulative sum of all previous measurements. This serves to consolidate measurements that agree, while allowing incremental updating with new data.&lt;/li&gt;
  &lt;li&gt;It allows us to model camera specific measurement uncertainty. That is, how a particular sensing technology (such as Time-of-Flight) affects depth measurement. The reliability of depth estimation for instance, might drop as a function of the distance from the camera’s optical centre (due the effects of lens distortion etc). A way to overcome this could be to design the weighting function to be a 2D Gaussian centered at the optical center.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Although certain sensing modalities might be more susceptible to unmodelled camera-specific measurement uncertainty, in practice, you can get reasonable results even if you ignore the second point - &lt;em&gt;i.e.&lt;/em&gt; assume &lt;script type=&quot;math/tex&quot;&gt;w_i(\mathbf{x})&lt;/script&gt; to be uniformly distributed (equal to &lt;script type=&quot;math/tex&quot;&gt;1&lt;/script&gt; at each step).&lt;/p&gt;

&lt;p&gt;Equations &lt;script type=&quot;math/tex&quot;&gt;\eqref{eq:eight}&lt;/script&gt; and &lt;script type=&quot;math/tex&quot;&gt;\eqref{eq:nine}&lt;/script&gt; give the update rules for the cumulative signed-distance and weight functions for each new frame (time step &lt;script type=&quot;math/tex&quot;&gt;i+1&lt;/script&gt;).&lt;/p&gt;

&lt;p&gt;That’s it! Combining all the steps mentioned so far allows us to compute a 3D reconstructed point-cloud. The pseudo-code below summarises the algorithm.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt; 
&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;idx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;voxel&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d_v&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;enumerate&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;grid&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;distances&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;voxel_in_frame&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;idx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;old_weight&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;weight_array&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;idx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;new_weight&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;old_weight&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;     &lt;span class=&quot;c&quot;&gt;# uniform weight distribution&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;d&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;min&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d_v&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;epsilon&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;       &lt;span class=&quot;c&quot;&gt;# truncated distance to surface&lt;/span&gt;

    &lt;span class=&quot;c&quot;&gt;# Update&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;tsdf_array&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;idx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;tsdf_array&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;idx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;old_weight&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;new_weight&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;weight_array&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;idx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;new_weight&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;The resulting point-cloud can be turned into a mesh like the one in &lt;em&gt;Figure 13&lt;/em&gt; using the &lt;a href=&quot;http://paulbourke.net/geometry/polygonise/&quot;&gt;marching-cubes algorithm&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/3D_TSDF.png&quot; alt=&quot;&quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 8: 3D mesh reconstructed with TSDF-Fusion.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Practical Considerations:&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Implementing a discrete voxel-grid based method is likely to be quite slow and require a GPU. For faster approaches you can look into &lt;a href=&quot;https://www.nvidia.com/object/nvidia_research_pub_018.html&quot;&gt;sparse voxel octrees&lt;/a&gt;.&lt;/li&gt;
  &lt;li&gt;The quality of the reconstruction greatly depends on your extrinsics calibration, and the viewing angle relative to a surface. The noise covariance ellipses of most depth cameras expand around regions tangential to viewing direction.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;experiments-with-depth-images&quot;&gt;Experiments with Depth Images&lt;/h2&gt;

&lt;p&gt;Sometimes it might be necessary to perform warping operations (resize, rotate, distort) on depth images where, unlike RGB images, the pixel values have a spatial meaning. An interpolation operation between adjacent pixels therefore cannot simply be expressed as a bi-linear (or bi-cubic, &lt;em&gt;etc.&lt;/em&gt;) average.&lt;/p&gt;

&lt;p&gt;This becomes problematic when interpolating around regions with dead pixels (pixels which have no corresponding depth measurement) or object edges. Dead pixels are common in almost every type of depth sensing modality; they appear around edges due to occlusions in stereo-imaging, on objects that are black in IR imaging (since black is a good sink for IR radiation), &lt;em&gt;etc.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Since depth images are actually just point-clouds, the common way of dealing with them is by performing operations in 3D. Libraries such as &lt;a href=&quot;http://www.pointclouds.org/&quot;&gt;PCL&lt;/a&gt; provide a number of convienient methods, such &lt;a href=&quot;http://docs.pointclouds.org/1.7.0/classpcl_1_1_bilateral_upsampling.html#details&quot;&gt;bilateral upsampling&lt;/a&gt;, to do such operations.&lt;/p&gt;

&lt;p&gt;Here we’ll consider an alternative approach by working with the “projection” of the point-cloud onto a 2D image. In particular, this approach is useful when generalising image pre-processing to depth images in a deep-neural-network pipeline. Some advantages of doing this are:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;You can operate directly on the 2D depth image, rather than converting it to a point-cloud format.&lt;/li&gt;
  &lt;li&gt;It allows us to perform arbitrary warping operations on depth images using pipelines designed for mono-images.&lt;/li&gt;
  &lt;li&gt;It allows us to construct a minimal solution that only addresses the interpolation part and leaves the mapping/filtering/smothing operation up to the user.&lt;/li&gt;
  &lt;li&gt;Given a 2D image, you can design an algorithm with &lt;script type=&quot;math/tex&quot;&gt;O(N)&lt;/script&gt; complexity, where &lt;script type=&quot;math/tex&quot;&gt;N&lt;/script&gt; is the output image size.&lt;/li&gt;
  &lt;li&gt;It is trivial to parallelise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;N.B.&lt;/em&gt;&lt;/strong&gt; If you’re doing conventional 3D computer vision, you’re probably better off in the long-run to manage point-cloud operations in 3D with PCL &lt;em&gt;etc&lt;/em&gt;. The method mentioned below is particularly designed for a deep-learning setting where you might want image warping as an augmentation technique, although you still need to keep track of it (since it affects the 3D position of points).&lt;/p&gt;

&lt;p&gt;Since the image pixels still have spatial meaning though, using out-of-the-box interpolation techniques developed for color images doesn’t work and leads to undesired “flying pixels” in the processed image and 3D-reconstruction &lt;em&gt;Figure 9&lt;/em&gt;.&lt;/p&gt;

&lt;div&gt;
&lt;table class=&quot;table_align&quot; style=&quot;width:100%&quot;&gt;
    &lt;tr&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/int_linear.png&quot; alt=&quot;Bilinear Interpolation&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(a)&lt;/i&gt;
        &lt;/td&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/linear_recon.png&quot; alt=&quot;Linear Reconstruction&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(b)&lt;/i&gt;
        &lt;/td&gt;
    &lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 9: (a) Naive bi-linear interpolation on a depth image results is flying pixels that look like fuzzy edges in 2D. (b) TSDF reconstruction makes the flying pixels more obviously visualisable.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Probably the easiest option is to use a nearest-neighbour interpolation, which should suffice for most cases, but performs somewhat poorly around edges, which become more jagged and imprecise, since you’re rounding off pixel coordinates. Additionally, noise present in the original also gets carried over to the warped image.&lt;/p&gt;

&lt;p&gt;Instead, let’s experiment with a custom algorithm for interpolating depth images, that:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Ignores regions with dead-pixels.&lt;/li&gt;
  &lt;li&gt;Elegantly handles interpolation around regions with object edges.&lt;/li&gt;
  &lt;li&gt;Smooths areas with outlier (noisy) pixel while maintaining edge boundaries (essentially a bilateral filter).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We’ll consider the case of bi-linear interpolation; a quadratic estimation of the value of fractional pixel coordinates &lt;script type=&quot;math/tex&quot;&gt;P&lt;/script&gt; based on the values of its four bounded-box pixels {&lt;script type=&quot;math/tex&quot;&gt;F_{00}, F_{01}, F_{10}, F_{11}&lt;/script&gt;}.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/bilinear_int.png&quot; alt=&quot;alt Bilinear Interpolation&quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 10: Bilinear Interpolation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;First, we need a mapping of pixels from the source (original) to the destination (warped) image. As an example, we can use the OpenCV &lt;code class=&quot;highlighter-rouge&quot;&gt;initUndistortRectifyMap&lt;/code&gt; function, which would normally be used to correct for lens distortion on monocular images. It returnes a handy map (&lt;code class=&quot;highlighter-rouge&quot;&gt;cmap&lt;/code&gt; in the code below) which tells us where each pixel in the destination image gets its value from in the source image, an operation that rarely yields integer pixel coordinates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;N.B.&lt;/em&gt;&lt;/strong&gt; You normally do image undistortion &lt;em&gt;before&lt;/em&gt; computing depth maps. Here we’re using it purely as an convenient example.&lt;/p&gt;

&lt;h4 id=&quot;ignoring-dead-pixels&quot;&gt;Ignoring Dead-pixels&lt;/h4&gt;

&lt;p&gt;Given &lt;code class=&quot;highlighter-rouge&quot;&gt;cmap&lt;/code&gt; we first ignore all pixel that lie outside the image boundaries. Next, we look up the value of each of the four bounding-box pixels. If two or more of these four pixels are zero, we consider this point to lie near a dead-pixel zone and assign a value of zero.&lt;/p&gt;

&lt;h4 id=&quot;edge-case-literally&quot;&gt;Edge Case (literally)&lt;/h4&gt;

&lt;p&gt;Next we consider regions that lie around object edges. To detect these, we compute the &lt;a href=&quot;https://en.wikipedia.org/wiki/Discrete_Laplace_operator&quot;&gt;&lt;em&gt;Laplacian&lt;/em&gt;&lt;/a&gt; of the source image (&lt;code class=&quot;highlighter-rouge&quot;&gt;cv::Laplacian&lt;/code&gt;). Technically the Laplacian is the divergence of the gradient at each pixel. On a discrete 2D grid it is approximated using a convolution with a 3 x 3 kernel:&lt;/p&gt;

&lt;script type=&quot;math/tex; mode=display&quot;&gt;% &lt;![CDATA[
\nabla^2 =
     \begin{bmatrix}
     0 &amp; 1 &amp; 0  \\
     1 &amp; -4 &amp; 1  \\
     0 &amp; 1 &amp; 0  \\
     \end{bmatrix} %]]&gt;&lt;/script&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;It acts by highlighting areas with rapid change in image gradients. This happens twice around each edge pixel; moving from a region with small gradients (non-edge pixels) to an area with high gradients (edge pixels) and back again (non-edge pixels). Depending on the relative pixel intensities on either side of an edge, the Laplacian yields two edges of pixels with opposite signs (&lt;em&gt;Figure 11a&lt;/em&gt;).&lt;/p&gt;

&lt;div&gt;
&lt;table class=&quot;table_align&quot; style=&quot;width:100%&quot;&gt;
    &lt;tr&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/laplacian.png&quot; alt=&quot;Edge Image&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(a)&lt;/i&gt;
        &lt;/td&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/edge_image.png&quot; alt=&quot;Laplacian Image&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(b)&lt;/i&gt;
        &lt;/td&gt;
    &lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 11:(a) Zoomed in image of an edge between two objects. (b) When moving across an edge, the Laplacian highlights two edges, one on each image. The signs and magnitudes of the pixel values at these edges depend on the relative pixel intensities when moving across the edge.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This doubly highlighted edge as result of the Laplacian allows us detect pixels belonging to objects on either side of the edge boundary and assign them to the correct one.&lt;/p&gt;

&lt;p&gt;To visualise this point better, consider the case in &lt;em&gt;Figure 12&lt;/em&gt;. It visualises the four possible cases that might arise when interpolating around edge pixels.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/edge_aware_concept.png&quot; alt=&quot;alt Edge Interpolation Concept &quot; /&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 12: Visualising interpolation around edge pixel belonging to two objects: A (Red) and B (Green). In each case, the interpolated point in the source image &lt;script type=&quot;math/tex&quot;&gt;p&lt;/script&gt; is marked with and x.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In &lt;em&gt;Figure 12a&lt;/em&gt;, three of the four bounding-box pixels lie on an edge. The top-left pixel &lt;script type=&quot;math/tex&quot;&gt;p_1&lt;/script&gt; belongs to the edge of object &lt;script type=&quot;math/tex&quot;&gt;A&lt;/script&gt; while the other two pixels belong to that of object &lt;script type=&quot;math/tex&quot;&gt;B&lt;/script&gt;. To detect this and assign the pixel to &lt;script type=&quot;math/tex&quot;&gt;B&lt;/script&gt;, we threshold the difference between the median values of the pixels and assign &lt;script type=&quot;math/tex&quot;&gt;p&lt;/script&gt; to the average of &lt;script type=&quot;math/tex&quot;&gt;p2&lt;/script&gt; and &lt;script type=&quot;math/tex&quot;&gt;p3&lt;/script&gt;. We can do this because the pixel values at {&lt;script type=&quot;math/tex&quot;&gt;p_1, p_2, p_3, p_4&lt;/script&gt;} have spatial meaning and therefore their difference indicates the distance between them. We can therefore decide for instance, that edge pixeles that are within 2 &lt;script type=&quot;math/tex&quot;&gt;cm&lt;/script&gt; of each other belong to the same object.&lt;/p&gt;

&lt;p&gt;A second edge case is where a horizontal or vertical edge boundary divides the bounding-box equally (&lt;em&gt;Figure 12c-d&lt;/em&gt;). In this case, we finally give up on interpolation (because there isn’t really a correct one!) and assign &lt;script type=&quot;math/tex&quot;&gt;p&lt;/script&gt; to its nearest neighbour.&lt;/p&gt;

&lt;p&gt;The code for implementing all of this is given below.&lt;/p&gt;

&lt;figure class=&quot;highlight&quot;&gt;&lt;pre&gt;&lt;code class=&quot;language-c--&quot; data-lang=&quot;c++&quot;&gt; 
&lt;span class=&quot;k&quot;&gt;using&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;namespace&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;std&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;using&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;namespace&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cv&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;

&lt;span class=&quot;kt&quot;&gt;void&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;edgeAwareRemap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;InputArray&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;OutputArray&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;InputArray&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_cmap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
    &lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;fxy&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;vector&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;bool&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;vector&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;Mat_&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;src&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;getMat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;getMat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;();&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;Mat&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cmap&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_cmap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;getMat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;();&lt;/span&gt;

    &lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;h&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rows&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;w&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cols&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;Mat_&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;Mat_&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;yMat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;CV_32F&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;xMat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;CV_32F&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;F&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;CV_32F&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;

    &lt;span class=&quot;n&quot;&gt;Laplacian&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;CV_32F&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;100&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;i&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;i&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;h&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;++&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;w&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;++&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;

            &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;clear&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;();&lt;/span&gt;

            &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cmap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Point2f&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;cmap&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Point2f&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;

            &lt;span class=&quot;c1&quot;&gt;// Ignore invalid indices
&lt;/span&gt;            &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;((((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;h&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;(((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;w&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mf&quot;&gt;0.0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
            &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;floor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;floor&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ceil&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ceil&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;

                &lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;at&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;

                &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;push_back&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt;
                    &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;push_back&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt;
                    &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;push_back&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt;
                    &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;push_back&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;||&lt;/span&gt;
                    &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;laplacianImage&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;edgeThreshold&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;));&lt;/span&gt;

                &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;any_of&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;begin&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;isEdgePx&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;end&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[](&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;bool&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;v&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;v&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;}))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;};&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;sort&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;begin&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;end&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;());&lt;/span&gt;

                    &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;20&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                        &lt;span class=&quot;c1&quot;&gt;// Assign median value
&lt;/span&gt;                        &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pxList&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;])&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
                        &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                    &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
                    &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                        &lt;span class=&quot;c1&quot;&gt;// Find and assign to nearest neighbour
&lt;/span&gt;                        &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;src&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;round&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)),&lt;/span&gt;
                                        &lt;span class=&quot;k&quot;&gt;static_cast&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;int&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;round&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)));&lt;/span&gt;
                        &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                    &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;

                &lt;span class=&quot;c1&quot;&gt;// Tolerate at-most one zero value for corner pixels
&lt;/span&gt;                &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
                &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
                &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
                &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;/&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
                &lt;span class=&quot;k&quot;&gt;else&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;!&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;((&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;!=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)))&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
                    &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                    &lt;span class=&quot;k&quot;&gt;continue&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;

                &lt;span class=&quot;c1&quot;&gt;// Interpolate
&lt;/span&gt;                &lt;span class=&quot;n&quot;&gt;yMat&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y1&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;y0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;xMat&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x1&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;x0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;F&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f00&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f01&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f10&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f11&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;fxy&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;Mat_&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;&lt;/span&gt;&lt;span class=&quot;kt&quot;&gt;float&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;yMat&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;F&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;*&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;xMat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
                &lt;span class=&quot;n&quot;&gt;dst&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;j&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;kt&quot;&gt;uint16_t&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;fxy&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
            &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
    &lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;};&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;

&lt;p&gt;The results of applying such a remapping operation are visualised in (&lt;em&gt;Figure 13b&lt;/em&gt;). In comparison, the nearest neighbour interpolation (&lt;em&gt;Figure13c&lt;/em&gt;) smoothed with a median filter has pronounced distortions at the edges (&lt;em&gt;Figure13d&lt;/em&gt;) when compared to the edge-aware interpolation. The median filter also dialates already distorted edges, creating spurious data.&lt;/p&gt;

&lt;p&gt;Our custom interpolation on the other-hand strikes a balance between edge-preserving warping, smoothing and accurate interpolation.&lt;/p&gt;

&lt;div&gt;
&lt;table class=&quot;table_align&quot; style=&quot;width:100%&quot;&gt;
    &lt;tr&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/raw_depth.png&quot; alt=&quot;Raw Depth Image&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(a)&lt;/i&gt;
        &lt;/td&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/raw_eaw.png&quot; alt=&quot;Edge Aware Interpolation&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(b)&lt;/i&gt;
        &lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/Interpolated_NN.png&quot; alt=&quot;Nearest Neighbour&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(c)&lt;/i&gt;
        &lt;/td&gt;
        &lt;td class=&quot;table_align&quot;&gt;
            &lt;img src=&quot;assets/img/tsdf/diff_image_median.png&quot; alt=&quot;Diff Image&quot; /&gt; &lt;br /&gt;
             &lt;i&gt;(d)&lt;/i&gt;
        &lt;/td&gt;
    &lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 13: Comparison of nearest-neighbour (c) and the custom interpolation (b) techniques from a warped raw image (a). (d) Difference image highlighting edge discrepancies between (b) and (c)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is ofcourse not the only way you could do this. Here we’ve approximated the interpolation with a median/nearest-neighbour approach when the source pixel lies on a edge. You could probably do better by using some other kind of average. But unless you really care about those diminishing returns, after a point the difference will be imperceptible.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Practical Considerations:&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The effects of using this kind of approach will depend on the structure of the scene being imaged (small and rapidly varying features vs big relatively smooth ones).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;applications&quot;&gt;Applications&lt;/h2&gt;

&lt;p&gt;While there have been leaps of progress in 2D computer vision, 3D vision has lagged behind, partly due to a lack of tools to deal with its geometric nature. Yet, the world is 3D and there are obvious advantages in treating it as such.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://geometricdeeplearning.com/&quot;&gt;Geometric Deep Learning&lt;/a&gt; has a list of the state of the art methods in deep-learning on graphs, a few of which include some pretty impressive work on 3D vision. Here are a few examples on the exciting use of 3D vision:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Shape segmentation with &lt;a href=&quot;https://arxiv.org/pdf/1612.00606.pdf&quot;&gt;SyncSpecCNN&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/syncspeccnn.png&quot; alt=&quot;syncspeccnn&quot; /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Dense shape correspondence with &lt;a href=&quot;https://arxiv.org/pdf/1704.08686.pdf&quot;&gt;Deep Functional Maps&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/deep_functional_maps.png&quot; alt=&quot;deep_functional_maps&quot; /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Learning visual descriptors for object using &lt;a href=&quot;https://arxiv.org/pdf/1806.08756.pdf&quot;&gt;Dense Object Nets&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/dense_visual_descriptors.png&quot; alt=&quot;dense_visual_descriptors&quot; /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Directly learning Signed Distance Functions with &lt;a href=&quot;https://arxiv.org/pdf/1901.05103.pdf&quot;&gt;DeepSDF&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;img src=&quot;assets/img/tsdf/deepsdf.png&quot; alt=&quot;deepsdf&quot; /&gt;&lt;/p&gt;

&lt;p&gt;3D computer vision is an exciting and important field. It deserves at least as much attention as 2D vision if not more! I hope this blog post helps a few more people approach it :)&lt;/p&gt;</content><author><name>Noorvir Aulakh</name></author><summary type="html"></summary></entry></feed>