\documentclass[]{IEEEtran}
\usepackage{cite}
\usepackage[tight,footnotesize]{subfigure}
\usepackage[cmex10]{amsmath}
\usepackage{amsfonts}
\usepackage{url,epstopdf}
\usepackage{graphicx}
\usepackage{subfig}
\usepackage[cmex10]{amsmath}
\usepackage{amsfonts}
\usepackage{algorithm}
\usepackage[noend]{algpseudocode}
\usepackage{amsmath,bm}
\usepackage{nopageno}

\pagestyle{empty}
\thispagestyle{empty}


% correct bad hyphenation here
\hyphenation{op-tical net-works semi-conduc-tor}


\begin{document}
%


% paper title
% can use linebreaks \\ within to get better formatting as desired
% Do not put math or special symbols in the title.
\title{A Video Aided RF Localization Technique for the Wireless Capsule Endoscope (WCE) inside Small Intestine}


% author names and affiliations
% use a multiple column layout for up to three different
% affiliations
\author{Guanqun Bao,
        Liang Mi,
        and~Kaveh Pahlavan,~\IEEEmembership{Fellow,~IEEE}\\
\IEEEauthorblockA{Center for Wireless Information Network Studies\\
Worcester Polytechnic Institute\\
Worcester, MA, 01609, USA\\
Email:(gbao, lmi, kaveh)@wpi.edu\\}
}
% conference papers do not typically use \thanks and this command
% is locked out in conference mode. If really needed, such as for
% the acknowledgment of grants, issue a \IEEEoverridecommandlockouts
% after \documentclass

% for over three affiliations, or if they all won't fit within the width
% of the page, use this alternative format:
% 
%\author{\IEEEauthorblockN{Michael Shell\IEEEauthorrefmark{1},
%Homer Simpson\IEEEauthorrefmark{2},
%James Kirk\IEEEauthorrefmark{3}, 
%Montgomery Scott\IEEEauthorrefmark{3} and
%Eldon Tyrell\IEEEauthorrefmark{4}}
%\IEEEauthorblockA{\IEEEauthorrefmark{1}School of Electrical and Computer Engineering\\
%Georgia Institute of Technology,
%Atlanta, Georgia 30332--0250\\ Email: see http://www.michaelshell.org/contact.html}
%\IEEEauthorblockA{\IEEEauthorrefmark{2}Twentieth Century Fox, Springfield, USA\\
%Email: homer@thesimpsons.com}
%\IEEEauthorblockA{\IEEEauthorrefmark{3}Starfleet Academy, San Francisco, California 96678-2391\\
%Telephone: (800) 555--1212, Fax: (888) 555--1212}
%\IEEEauthorblockA{\IEEEauthorrefmark{4}Tyrell Inc., 123 Replicant Street, Los Angeles, California 90210--4321}}

% use for special paper notices
%\IEEEspecialpapernotice{(Invited Paper)}

% make the title area
\maketitle

\pagestyle{empty}
\thispagestyle{empty}


% As a general rule, do not put math, special symbols or citations
% in the abstract
\begin{abstract}
Wireless capsule endoscope (WCE) provides a noninvasive method to examine the entire gastrointestinal (GI) tract including small intestine, which other video endoscopic instruments cannot reach.  Since the shape of small intestine is extremely complex and the length of small intestine varies from 5 to 9 meters, localization of the WCE inside the small intestine is very challenging. Traditional radio frequency (RF) localization techniques using the received signal strength (RSS) are not able to provide satisfactory location information of the capsule inside the small intestine. In this paper, we present a hybrid localization technique that takes advantage of data fusion from image sequence captured by the WCE's embedded camera and the RSS of the RF signal emitted by the capsule to enhance the positioning accuracy. The proposed method estimates the speed and direction of movement of the capsule by analyzing displacements of feature points between consecutive image frames and this motion information is integrated with RSS measurements by employing a Kalman filter to smooth the RF localization results. Performance of the proposed method is validated under a virtual testbed that emulates the transition of capsule inside small intestine against the traditional RSS-based RF localization.
\end{abstract}

% no keywords
\begin{IEEEkeywords} Wireless capsule endoscope, hybrid localization \end{IEEEkeywords}.
% For peer review papers, you can put extra information on the cover
% page as needed:
% \ifCLASSOPTIONpeerreview
% \begin{center} \bfseries EDICS Category: 3-BBND \end{center}
% \fi
%
% For peerreview papers, this IEEEtran command inserts a page break and
% creates the second title. It will be ignored for other modes.
\IEEEpeerreviewmaketitle

\section{Introduction}
% no \IEEEPARstart
Wireless capsule endoscopy (WCE) \cite{faigel2008capsule} is progressively emerging as a popular non-invasive imaging tool for gastrointestinal (GI) tract diagnosis. Compared with the traditional colonscope or enterscope, WCE has the capability of examining the entire small intestine, which other endoscopic instruments can not reach. However, since the length of small intestine is too long (varies from 5m to 9m \cite{pahlavan2012rf}) and it is twisted inside the abdominal cavity with extremely indistinguishable distribution, the localization of the capsule inside small intestine becomes very challenging, which prevents physicians from administering immediate therapeutic operations after an abnormality is found by the video source. Thus, having a precise and reliable localization system for the capsule inside the small intestine would greatly enhance the benefits of WCE. During the past few years, many attempts have been made to develop accurate and reliable localization systems for the WCE \cite{de2009intestinal, salerno2012discrete}. A commonly used localization infrastructure, which has been chosen for commercial use for M2A capsule \cite{jacob2001localization} designed by Given Imaging, is to attach many calibrated external antennas to the anterior abdominal wall of the human body to detect the RF signal emitted by the wireless capsule \cite{ye2012accuracy}. By interpreting the power of the received signal into distance between the capsule and body mounted sensor array, position of the capsule can be estimated by pattern matching algorithms such as least square algorithm \cite{huang2001real} and maximum likelihood algorithm \cite{shenli2012}. However, due to the non-homogeneity and severe attenuation of body tissues, features of the received signal are sometimes poorly correlated with the distance. Therefore, this RF localization system often end up providing discontinuous and scattered estimations with unacceptable amount of error \cite{de2009intestinal}.

One way to enhance the performance of RF localization is to combine the motion information of the capsule by employing a data fusion algorithm such like Kalman filter or particle filter \cite{pahlavan2012rf}. In the localization literature, there has been a trend to extract motion parameters from  image sequence to improve the accuracy of RF localization. This class of algorithms is known as video based simultaneous localization and mapping (SLAM) algorithms \cite{silveira2008efficient}. In the WCE application, since the endoscopic capsule continually takes pictures with very short time interval (two frames / sec), it is possible to reconstruct motion information of the capsule from video stream \cite{baomotion}. In this paper, we present a hybrid localization technique that is able to extract speed and moving direction of the endoscopic capsule from endoscopic image frames to aid the RF localization. The major contribution of this paper is that we explored the potential of using images as another source to track the position of WCE and we established a virtual platform to validate our algorithm.

\begin{figure*}[!t]
\centering
\subfigure[Feature points matching between consecutive frames]{
   \includegraphics[scale =0.5] {Fig2a.eps}
   %\label{fig:Fig2a}
 }
 \subfigure[Motion vectors by linking corresponding feature points]{
   \includegraphics[scale =0.5] {Fig2b.eps}
   %\label{fig:Fig2b}
 }
\caption{Feature points used for motion detection}
\label{fig:feature}
\end{figure*}

The rest of the paper is organized as follows: In section II, we explain how to track the motion of the capsule based on the image sequence captured by the endoscopic camera. Section III describes how to integrate the motion information extracted from images with RF measurements using a Kalman filter to enhance the localization accuracy. In section IV, performance of the proposed hybrid localization algorithm is validated using a virtual testbed against the traditional RF localization algorithm. Finally, conclusion and future work are addressed in section V.


\section{Motion Tracking using Images}

Since the endoscopic capsule continuously takes pictures at two frames/sec as it travels, it's possible to obtain information such like how far the capsule has moved and the direction of movement by analyzing the displacements of unique portion of the scene, which referred to as feature points (FPs), between consecutive image frames. According to the literature \cite{lee2011motion, li2011motion} and our previous works \cite{baomodeling, Baoemulation}, more FPs can be accurately detected by the Affine Scale-invariant Feature Transform (ASIFT) algorithm \cite{morel2009asift} compared to other algorithms. An example of feature matching result is given in Fig.~\ref{fig:feature} (a). ``o'' represents the coordinates of detected FPs in the reference (first) frame, ``*'' represent the coordinates of matched FPs on the second frame. If we connect the corresponded FP pairs on the same frame (as shown in Fig.~\ref{fig:feature} (b)), a bunch of motion vectors will be generated representing the displacements of FPs between frames. The magnitudes and distribution of these motion vectors reflect the motions such as speed and direction of moving of the endoscopic capsule during the elapsed time interval.

\subsection{Image ``Unrolling'' for Motion Detection}

To standardize the displacement of each FP pair and facilitate the quantitive calculations of motion parameters that are useful for localization, we need to perform an inverse cylindrical projection \cite{tillo2010inverse} (also referred as ``image unrolling'' in \cite{sathyanarayana2007real}) (shown in Fig.~\ref{fig:imageunroll}) to project the original cylindrical image onto a flatten view coordinate system, which we called ``unrolled'' image domain. Given a FP $P$ at distance $d$ away from the camera, the angular depth of $P$ is defined as:

\begin{equation} \label{eq1}
 \theta= tan^{-1}\left(\frac{R}{d}\right)
\end{equation}
where $R$ represents the radius of the intestinal tube. It can be seen from Eq.~\ref{eq1} that a smaller angler depth indicates a larger distance away from the camera. To facilitate the derivation of angler depth, we map the coordinate $(x, y)$ of any point on the cylindrical image domain to the unrolled image domain $(x', y')$ by: 

\begin{figure}[!t]
\begin{center}
\begin{tabular}{c}
\scalebox{0.67}{\includegraphics[]{imageunroll.eps}}
\end{tabular}
\caption{The process of "unrolling" the cylindrical image}
\label{fig:imageunroll}
\end{center}
\end{figure}

\begin{equation} \label{eq2}
 x'=\frac{L\phi}{2\pi}  \quad \quad y'=r
\end{equation}
where $\phi$ is the angle between point $P$ and the horizontal axis in the cylindrical image domain (shown in Fig.~\ref{fig:imageunroll} (a)).
\begin{equation} \label{eq3}
 \phi =tan^{-1}\left(\frac{y-y_0}{x-x_0}\right)
\end{equation}
$r$ is the radius of the circular ring associated with point $P$ that can be calculated by:
\begin{equation} \label{eq4} 
 r= \sqrt {(x-x_0)^2+(y-y_0)^2}.
\end{equation}
$L$ and $H$ are the length and the height of the unrolled image domain respectively. $\eta$ is the field of view. Under this new coordinate system, the angular depth of any point $P$ can be calculated directly through its $y'$ value by:

\begin{equation} \label{eq5}
 \theta\cong\left(\frac{y'}{H}\right)\eta
\end{equation}


\begin{figure*}[!t]
\begin{center}
\begin{tabular}{c}
\scalebox{0.55}{\includegraphics[]{speed.eps}}
\end{tabular}
\caption{Geographic model for speed estimation}
\label{fig:speed}
\end{center}
\end{figure*}


%\begin{figure*}[!t]
%\begin{center}
%\begin{tabular}{c}
%\scalebox{0.65}{\includegraphics[]{speed.eps}}
%\end{tabular}
%\caption{Speed estimation of the video capsule }
%\label{fig:speed}
%\end{center}
%\end{figure*}

\subsection{Quantitive Calculation of Speed and Direction of Motion}

\subsubsection{Estimation on Speed}

As mentioned in previous sections, motions of a video capsule can be detected by measuring the displacements of the FPs. To explain better, we use Fig.~\ref{fig:speed} to illustrate the procedure of calculating the transition speed of a capsule traveling through the intestinal tube. Point $P$ is a FP detected at a distance $D$ from the initial position of the camera $C$ with its angular depth equals to $\theta_1$. After the camera has moved forward by a distance $d$ to a new position $C'$, the angular depth of $P$ changes to $\theta_2$. The changes in angular depth can be used to calculate the transition speed of the capsule. 

\begin{equation} \label{eq6}
 \theta_1=tan^{-1}\frac{R}{D} \quad \Longrightarrow \quad D=\frac{R}{tan\theta_1}
\end{equation}

\begin{equation} \label{eq7}
 \theta_2=tan^{-1}\frac{R}{D-d}
\end{equation}
Replacing $D$ in Eq.~\ref{eq7} with Eq.~\ref{eq6}, we get:

\begin{equation} \label{eq8}
 d=\frac{R}{tan\theta_2}\left(1-\frac{tan\theta_2}{tan\theta_1}\right)
\end{equation}
since the time interval for this distance $d$ is half a second, the speed of the capsule can be calculated by:

\begin{equation} \label{eq9}
 v=\frac{\frac{1}{N}\sum_{i=0}^N d_i}{0.5}=\frac{2}{N}\sum_{i=0}^N\frac{R}{tan\theta_{2i}}\left(1-\frac{tan\theta_{2i}}{tan\theta_{1i}}\right)
\end{equation}
where $N$ equals to the total number of all detected FPs. To reemphasize, the unrolling process facilitates the deriving of $\theta_1$ and $\theta_2$ in Eq.~\ref{eq5} and therefore facilitates the deduction of $v$. Similarly, if the capsule moves backward, the speed can be calculated in the same manner as well.

\subsubsection{Estimation on direction of moving}

\begin{figure}[!t]
\begin{center}
\begin{tabular}{c}
\scalebox{0.5}{\includegraphics[]{motiontracking.eps}}
\end{tabular}
\caption{Direction of moving of the capsule}
\label{fig:motiontracking}
\end{center}
\end{figure}


Another important aspect for motion tracking is estimating the direction of moving of the capsule. If we define the world coordinate as $(X, Y, Z)$ and capsule's coordinate as $(X', Y', Z')$. The moving direction of the capsule is given by a norm vector $(n_x, n_y, n_z)^T$ in the world coordinate.  After the capsule rotated with angle $\alpha$ around its $X'$ axis (pitch), angle $\beta$ around its $Y'$ axis (yaw) and angle $\gamma$ around its $Z'$ axis (roll), the new direction of the capsule $(n_x', n_y', n_z')^T$ can be calculated by:

\begin{equation} \label{eq10}
\begin{bmatrix}n_x' \\n_y' \\n_z' \end{bmatrix}=\mathbb{R}\cdot\begin{bmatrix}n_x \\n_y \\n_z \end{bmatrix}
\end{equation}
where $\mathbb{R}$ is an accumulative rotation matrix which relates the camera's coordinate system $(X', Y', Z')$ to the world coordinate system $(X, Y, Z)$. If we assume the the camera's coordinate system was initially aligned with the world coordinate system with it's focal axis pointed to the $Z$ axis, then, the initial value of $\mathbb{R}$ equals to a $3\times3$ identical matrix. As the capsule moves away from the original position, $\mathbb{R}$ is updated at each time step by: 

\begin{figure}[!t]
\centering
\subfigure[Estimation of $\alpha$ and $\beta$]{
   \includegraphics[scale =0.5] {tiltangle.eps}
   %\label{fig:Fig2a}
 }
 
 \subfigure[Estimation of $\gamma$ ]{
   \includegraphics[scale =0.5] {rotation.eps}
   %\label{fig:Fig2b}
 }
\caption{Feature points used for motion detection}
\label{fig:angleestimation}
\end{figure}

\begin{equation} \label{eq11}
\mathbb{R}=\mathbb{R}\cdot\mathbb{R}_t\cdot\mathbb{R}^{-1}
\end{equation}
where $\mathbb{R}_t$ is an direction updating rotaton matrix that has a following expression:\\\\
$\mathbb{R}_t=$\\\\
$\scriptsize{\begin{bmatrix}cos\alpha cos\gamma & cos\gamma sin\alpha sin\beta - cos\alpha sin\gamma & cos\alpha cos\gamma sin\beta-sin\alpha sin\gamma \\cos\beta sin\gamma & cos\alpha cos\gamma + sin\alpha sin\beta sin\gamma & -cos\gamma sin\alpha + cos\alpha sin\beta sin\gamma \\ -sin\beta & cos\beta sin\alpha & cos\alpha cos\beta \end{bmatrix}}$
\begin{equation} \label{eq12}
\end{equation}
where $\alpha$, $\beta$ and $\gamma$ are the pitch, yaw and roll angles about the capsule's $X'$, $Y'$ and $Z'$ axises respectively during the elapsed time interval. Again, these angles can be obtained without complicated computation in the unrolled image domain.



\begin{itemize}
	\item pitch ($\alpha$) and yaw ($\beta$) estimation
\end{itemize}

As illustrated in Fig.~\ref{fig:angleestimation} (a), point $P$ and point $Q$ are of the same distance from the initial position of the camera $C$. After the camera tilting with angle $\varphi$ towards $Q$, the angular depths of the two FPs change with different amounts of magnitude.  The magnitude of tilting can be estimated by:

\begin{equation} \label{eq14}
 \varphi \cong \frac{\triangle Q-\triangle P}{max(\triangle P, \triangle Q)} 
\end{equation}

The direction of tilting can be obtained by finding the group with smallest displacement in $y'$ in the unrolled image domain. Therefore, this tilting angle $\phi$ can be further decomposed into pitch angle $\alpha$ and yaw angle $\beta$ by:

\begin{equation} \label{eq15}
 \alpha = \varphi \cdot cos\phi \quad  \beta = \varphi \cdot sin\phi
\end{equation}


\begin{itemize}
	\item roll ($\gamma$) estimation
\end{itemize}

The calculation of roll angle $\gamma$ is even easier in the unrolled image domain by measuring the horizontal displacements of FPs on the $x'$ axis (as shown in Fig.~\ref{fig:angleestimation} (b)):

\begin{equation} \label{eq16}
 \gamma=\frac{1}{N}\sum_{i=0}^N\frac{\triangle x_i'}{L}2\pi
\end{equation}

where $\triangle x'$ denotes the horizontal displacement of a FP in the unrolled domain. $L$ is the length of the unrolled image. 

%For estimating rotations, there is no need for calculating angular depths. This is because irrespective of how deep the points are in the cylinder, the apparent rotations remain the same. 

%\begin{figure*}
%\begin{center}
%\begin{tabular}{c}
%\scalebox{0.55}{\includegraphics[]{Kalman.eps}}
%\end{tabular}
%\caption{A complete picture of the operation of the Kalman filter}
%\label{fig:flowchart}
%\end{center}
%\end{figure*}

\section{Integration of Motion Tracking from Images with RF Signal}

In this section, we talk about how to use a Kalman filter to fuse the data from both sensors to improve the reliability and accuracy for determining the position of a video capsule inside the human body. 

Given the motion parameters derived from the last section, the priori motion state ${\widehat{\bm{m}}^{-}}_t$ at time step $t$ (without any knowledge of RF measurement) is given by:

\begin{equation} \label{eq17}
{\widehat{\bm{m}}^{-}}_t=\bm{A_{t-1} \cdot {\widehat{\bm{m}}}_{t-1}+\bm{\omega_{t-1}}}
\end{equation}
motion state vector $\bm{m_t}$ is defined as $\bm{m_t}=[x,y,z,n_x,n_y,n_z]^T$, where $x,y,z$ are the coordinates of the capsule in the world coordinate system and $[n_x,n_y,n_z]$ is a norm vector that indicates the direction of moving of the capsule. $\bm{\omega_t}$ is a noise term caused by inaccurate motion estimation which follows a normal probability distributions with covariance equal to $\bm{Q}$. $\bm{A}$ is a $6 \times 6$ transition matrix that relates the previous motion state at time $t-1$ to the current motion state at time $t$. If we plug in all the parameters, Eq.~\ref{eq17} can be rewritten as:  

\begin{equation} \label{eq18}
\begin{bmatrix}x_{t} \\y_{t} \\z_{t}\\n_{x|t}\\n_{y|t}\\n_{z|t} \end{bmatrix}=\begin{bmatrix}1 & 0 & 0 & v \triangle t & 0 & 0 \\ 0 & 1 & 0 & 0 & v \triangle t & 0 \\ 0 & 0 & 1 & 0 & 0 & v \triangle t \\ 0 & 0 & 0 &  &  &  \\ 0 & 0 & 0 &  &  \begin{bmatrix} \mathbb{R} \end{bmatrix}  &  &\\ 0 & 0 & 0 &  &  &  \end{bmatrix} \cdot\begin{bmatrix}x_{t-1} \\y_{t-1} \\z_{t-1}\\n_{x|t-1}\\n_{y|t-1}\\n_{z|t-1} \end{bmatrix} 
\end{equation}
where $v$ is the transition speed of the capsule derived from Eq.~\ref{eq9}. $\triangle t$ is the time interval between frames.  $\mathbb{R}$ is the same rotation matrix introduced in Eq.~\ref{eq10}. 

We can use the motion state to predict the upcoming RF localization $\bm{\widehat{z}_t}$ by: 

\begin{equation} \label{eq19}
\bm{\widehat{z}_t}=\bm{H_t}\cdot\bm{\widehat{m}^{-}_t}+ \bm{\nu_t} 
\end{equation}
where $\bm{\nu_t}$ is a measurement-obtained noise term. Similar to $\bm{\omega_t}$, $\bm{\nu_t}$ also followed a normal distribution with covariance equal to $\bm{R}$. $\bm{H}$ is a $3 \times 6$ matrix which relates the RF localization to the priori motion state at time $t$.
\begin{equation} \label{eq20}
\bm{H}=\begin{bmatrix}1 & 0 & 0 & 0 & 0 & 0\\0 & 1 & 0 & 0 & 0 & 0 \\ 0 & 0 & 1 & 0 & 0 & 0 \end{bmatrix} 
\end{equation}

The actual RF localization results are obtained using a least square algorithm introduced in \cite{ye2012accuracy}. Once the actual RF localization result $\bm{z_t}$ is available, we can use the priori motion estimate ${\widehat{\bm{m}}^{-}}_t$ and a weighted difference between the actual RF measurement $\bm{z_t}$ and the predicted RF measurement $\bm{\widehat{z}_t}$ to correct the localization results.

%\begin{figure}[!t]
%\begin{center}
%\begin{tabular}{c}
%\scalebox{0.35}{\includegraphics[]{Fig1.eps}}
%\end{tabular}
%\caption{A typical RF localization infrastructure for VCE.( d1, d2, and d3 are the RSS ranging distances between the capsule and three body mounted sensors respectively. Using $d1$ $d2$ and $d3$ to draw red circles around these sensors, the intersection of these circles should be the position of the capsule.)}
%\label{fig:RFlocalization}
%\end{center}
%\end{figure}

\begin{equation} \label{eq23}
{\widehat{\bm{m}}}_t={\widehat{\bm{m}}^{-}}_t+\bm{K}_t\left(\bm{z}_t-\bm{\widehat{z}_t}\right)
\end{equation}
where ${\widehat{\bm{m}}}_t$ is defined as a posteriori motion state estimate given the RF measurement $\bm{z}_t$. The $3 \times 6$ matrix $\bm{K}$ in Eq.~\ref{eq23} is called Kalman gain. If we define the priori estimate errors covariance as $\bm{P}^{-}_t=E[(\bm{m}_t-\widehat{\bm{m}}^{-}_t)(\bm{m}_t-\widehat{\bm{m}}^{-}_t)^T]$ and a posteriori estimate errors covariance as $\bm{P}_t=E[(\bm{m}_t-\widehat{\bm{m}}_t)(\bm{m}_t-\widehat{\bm{m}}_t)^T]$, the Kalman Gain can be expressed as:

\begin{equation} \label{}
\bm{K}_t=\bm{P}^{-}_t \bm{H}^T\left( \bm{H}\bm{P}^{-}_t\bm{H}^T +\bm{R}\right)^{-1}
\end{equation}
The Kalman Gain control the weighs of both sensors on the final position estimation: if RF measurement noise is low, then the final estimation is more dependent on the RF measurement. Otherwise, the final estimation is more dependent on the motion model. 

\section{Results and Discussion}

\begin{figure}[!t]
\begin{center}
\begin{tabular}{c}
\scalebox{0.47}{\includegraphics[]{emulation.eps}}
\end{tabular}
\caption{Emulation testbed set up}
\label{fig:emulation}
\end{center}
\end{figure}

\begin{figure}[!t]
\centering
\subfigure[Localization results of different algorithms]{
   \includegraphics[scale =0.45] {patha.eps}
   %\label{fig:Fig2a}
 }
 
 \subfigure[Evolution of localizatoin error as the capsule moves ]{
   \includegraphics[scale =0.5] {pathb.eps}
   %\label{fig:Fig2b}
 }
\caption{Localization results of different algorithms and performance evaluation}
\label{fig:result}
\end{figure}

\begin{figure}[!t]
\begin{center}
\begin{tabular}{c}
\scalebox{0.42}{\includegraphics[]{CDF.eps}}
\end{tabular}
\caption{Performance evaluation by CDF plot of different algorithms}
\label{fig:CDF}
\end{center}
\end{figure}

One of the major difficulties of implementing any capsule localization algorithm when it comes inside the human body is validation. That's because we have limited control of the capsule after it is swallowed by the patient so we could not verify the performance of the algorithms \cite{pahlavan2012rf, france2005layered}. Besides, carrying out experiments on the real human beings is extremely costly and restricted by law. Thus, the only way to test our localization algorithm is to build up an emulation test bed. To create a similar scenario, we generated a cylindrical tube with the same size and shape in a virtual 3D space. As illustrated on the top of Fig.~\ref{fig:emulation} (c), the virtual test bed shared the same topology with the real small intestine which is intertwined back and forth. To make the interior of the test bed look more realistic, we extracted color and texture from the real endoscopic images and mapped it onto the interior surface of the tube. The transition of the capsule was emulated by moving a virtual camera view point along the cylindrical tube. In this way, the movement of the camera can be fully controlled and the performance of the motion tracking algorithm can be validated. Also, to create a similar illumination effect of the real endoscopic image shown in Fig.~\ref{fig:emulation} (d), we placed a Phong light source behind the camera view point to emulate the LED light around the camera. Similar emulation set up can be found in \cite{france2005layered, szczypinski2009model}. 


%However, when using the video based tracking alone, the motion estimation error accumulates so that the estimated positions would gradually drift away from the actual positions as the capsule moves along. On the other hand, although the RF localization suffers great errors due to extended channel fading, it provides relatively more consistent localization accuracy which is independent from the previous measurements. A better solution is to take advantage of natures of both sensors and properly weigh the data according to their reliability. Kalman filter, which is a set of equations that implement self-corrective estimator, is well known as a powerful tool for data fusion. In this section, we present a hybrid technique for accurate localization of the endoscopy capsule inside small intestine by integrating data from video sensor and RF sensors using a Kalman filter. 

The results of motion tracking using images, RF localization and the proposed hybrid localization are given in Fig.~\ref{fig:result}. It can be seen from Fig.~\ref{fig:result} (a) that the results of RF localization (represented in green traigles) are scattered all around the small intestine with relative large error. This is because the RF channel suffers shadow fading and non-homogeneity of the body tissues. However, the good part of RF localization is its independent characteristics. Each measurement is an isolated procedure which cannot be affected by the previous measurements. Therefore, the localization error would not accumulate as the capsule moves along (shown in Fig.~\ref{fig:result}(b)). 

The result of the camera motion tracking algorithm is shown in gray line in Fig.~\ref{fig:result (a)} .  It shows that when using this algorithm alone, the estimated positions are continuous and the overall trend of the trajectory matches the ground truth path (shown in red line in Fig. ~\ref{fig:result} (a)) of the small intestine. However, as the capsule moves along, the localization error increases.  It can be seen in Fig.~\ref{fig:result} (b), after about 15 steps, the localization errors reaches to almost the same level of RF localization algorithm and it keeps increasing until explode. That's because the camera motion tracking is not an independent procedure, each estimated position is dependent on the previous estimations, thus, the error accumulated.  As the capsule moves further away from the beginning point, the trajectory begins to drift away from the correct path. Nevertheless, the transformation between every consecutive step is accurate, thus, when using this algorithm, the overall trend of the movement is still reliable, and it can provide continuous estimation of the capsule's movement. 

Finally we evaluated the performance of our proposed hybrid localization algorithm. The results are shown in purple line in Fig.~\ref{fig:result} (a). It shows that the localization results of hybrid localization are continuous and match the ground truth path of the small intestine very well. From Fig.~\ref{fig:result} (b) we can see that, compared with the previous two algorithms, the localization error of hybrid localization stays stable at a very low level and the error would not increase as the capsule moves along. The Cumulative error distribution function (CDF) of three localization algorithms are shown in Fig.~\ref{fig:CDF}.

\section{Conclusion}
In this paper, we presented a hybrid localization technique that utilizes camera motion tracking algorithm to aid the existing RF localization infrastructure for the WCE application. The performance of the proposed method is validated under a virtual emulation environment. The major contribution of this work is that we demonstrated the potential of using video source to aid the RF localization of the WCE. The proposed motion tracking technique is purely based on the image sequence that captured by the video camera which is already equipped on the capsule, thus, no extra components such as Inertial measurement units (IMUs) or magnetic coils are needed. Experimental results show that by combining the motion information with RF measurements, the proposed hybrid localization algorithm is able to provide accurate, smooth and continuous localization results. In the future, we will focus on refining this algorithm according to the clinical data and testing this algorithm with real human subjects. 

% conference papers do not normally have an appendix

% use section* for acknowledgement
\section*{Acknowledgment}

The authors would like to thank Dr. David Cave at UMass Memorial Medical Center for his precious suggestions, and the colleagues at the CWINS laboratory for their directly or indirectly help in preparation of the results presented in this paper.





% trigger a \newpage just before the given reference
% number - used to balance the columns on the last page
% adjust value as needed - may need to be readjusted if
% the document is modified later
%\IEEEtriggeratref{8}
% The "triggered" command can be changed if desired:
%\IEEEtriggercmd{\enlargethispage{-5in}}

% references section

% can use a bibliography generated by BibTeX as a .bbl file
% BibTeX documentation can be easily obtained at:
% http://www.ctan.org/tex-archive/biblio/bibtex/contrib/doc/
% The IEEEtran BibTeX style support page is at:
% http://www.michaelshell.org/tex/ieeetran/bibtex/
%\bibliographystyle{IEEEtran}
% argument is your BibTeX string definitions and bibliography database(s)
%\bibliography{IEEEabrv,../bib/paper}
%
% <OR> manually copy in the resultant .bbl file
% set second argument of \begin to the number of references
% (used to reserve space for the reference number labels box)
%\begin{thebibliography}{1}
%
%\bibitem{IEEEhowto:kopka}
%H.~Kopka and P.~W. Daly, \emph{A Guide to \LaTeX}, 3rd~ed.\hskip 1em plus
%  0.5em minus 0.4em\relax Harlow, England: Addison-Wesley, 1999.
%
%\end{thebibliography}

\ifCLASSOPTIONcaptionsoff
  \newpage
\fi


\bibliographystyle{IEEEtran}
\bibliography{reference}


% that's all folks
\end{document}


