Search      Print this chapter      Cite this chapter

IMAGE PROCESSING

Douglas G. Myers

Department of Computer Systems Engineering, School of Electrical and Computer Engineering, Curtin University of Technology, Perth, Western Australia

Keywords: Digital image processing, human visual response, image perception, image processing system, sampling, interpolation, image filtering, re-sampling. Feature extraction, pattern recognition, image understanding, edge detection

Contents

1. Introduction

2. Some Comments on Vision

3. What Is An Image?

4. The Relationship between Digital And Analog Images

5. The Concept of An Image Processing System

6. The Process of Image Formation

7. The Image as A Representation

8. The Image Processing Hierarchy

9. The Pre-Processing Level

10. Low Level Image Processing

11. Medium Level Image Processing

12. Image Interpretation

13. Interpolation in Image Processing

14. The Edge Detection Problem

15. Applications of Image Processing

16. Some Image Processing Packages

Related Chapters

Glossary

Bibliography

Biographical Sketch

Summary

Image processing, or more specifically digital image processing is one of the many important specialist areas of computing. It is an enabling technology for a wide range of applications including remote sensing, security, image databases, digital television and robotics. This chapter reviews the hierarchical four level structure of image processing techniques. Pre-processing techniques are designed to remove distortions introduced by sensors. Low level image processing techniques are mathematical or logical operators that perform simple processing tasks. Medium level image processing combines the simple low level operators to perform feature extraction and pattern recognition functions. High level image processing uses combinations of medium level functions to perform interpretation. Since this last level is effectively modeling the human visual response, the chapter briefly notes features of vision and perception. It also highlights two key image processing actions. Interpolation is a fundamental geometric operation used, for example, to map images to geographical coordinates. Edge detection is an operation that can be simple or, as when at the heart of many recognition tasks, quite complex. It serves as an excellent illustration of the distinctions in the image processing hierarchy. The chapter concludes with a brief examination of some of the key application areas and some commonly available image processing packages.

1. Introduction   

In the early 19th century, the development of photography provided a mechanical means of capturing what the human eye could see. Individuals could now record their likeness and, instead of reading about far away places, see a detailed representation of them. By the late 19th century, photography was both a popular pastime and an integral part of newspapers, books, journals and magazines. Then came the ‘moving picture’ and by the 1920s, a major new entertainment medium. In the 1930s, color photography was developed and the first television service began in Britain. World War II saw the rise of radar, with the location of detected objects being displayed as an electronic image. The first color television service began in the United States in January 1954 and also about that time the broadcast videotape recorder was created. These developments had related scientific and industrial outgrowths such as photographic film to record X rays or infrared radiation. Some innovative image capture systems were also created such as Schlieren photography. Originally developed to detect flaws in glass, it is probably best recognized now for showing the shock waves or turbulence about aircraft wings.

In spite of the increasing sophistication of the technology, in a strict technical sense all that had been developed until the 1950s was systems for image capture and storage. No system existed where, for example, images taken of microscope slides could be automatically processed to detect cancer cells.

The development of the laser in the 1950s provided a means of achieving that. The coherency of laser light allows a system to be set up to detect particular objects within photographs as well as other processing actions. As these systems require an optical bench with precision optical components such as lenses, gratings and filters, the number of applications for this approach has been limited. One of the first was processing the output of synthetic aperture radars (SARs), or sideways radars, used in remotely sensing terrain for exploration and mapping.

The early space probes could sense the physical properties of the planets, but the natural human inclination was to view them. Thus image sensors were installed, the captured images were digitized, compressed to reduce the data volume and then encoded so they could be communicated accurately bit by bit back to earth over the very long distances involved. On reception, the data was passed to a computer to check the coding and decompress the image. The sensors, though, distorted the images in various ways, particularly geometrically. Given a computer can perform a complex sequence of mathematical operations on any input; then it was natural to program the computer to correct this distortion.

This describes a system where an image is an input to a processing system. Thus image processing as we now largely know it began. That is to say, digital image processing. The cost of computing in the 1970s limited the field, but as that cost fell in successive decades, the impact of digital image processing began to rise significantly. Now, almost all images routinely seen in newspapers, books, advertising displays, magazines, journals and on television have been processed to some extent digitally. Further, computing has now reached a point where any user of a personal computer can acquire a range of software products at a reasonable price to undertake quite sophisticated image processing of almost any form. Very often, this is on images captured by that user’s own digital camera.

Image processing may be categorized in several ways. In terms of the means by which it is implemented, there is analog and digital image processing, but here only the latter will be considered. In terms of the broad focus, it may be described as analytic or synthetic. Synthesizing (digital) images is part of computer graphics and will also not be considered. Nevertheless, it is important to mention that the boundaries between analysis and synthesis can be blurred. For example, many special effects in film, television and advertising include elements of analysis as well as synthesis. To illustrate, in the Lord of the Rings films, scenes were needed of a dark environment with an active volcano. To achieve it, film was shot of an extinct volcano in New Zealand, it was processed to give the environment and the eruption was synthesized. Some modern areas of medical image processing such as virtual surgery also combine elements of analysis and synthesis.

In terms of the focus of analysis, image processing may be described as objective or subjective. Objective focuses on the functional. For example, it may be to detect particular objects within the image or to transform it in a particular way such as by removing noise or generating false colors. The outcome of processing, therefore, can be an image but it may also be more mundane such as a description of the number of each particular object identified and their properties. Subjective means the focus is essentially aesthetic; to manipulate the image in some way to improve its appeal to a human viewer. Here, the outcome is always an image. The steps leading to it, though, will involve objective actions such as filtering.

2. Some Comments on Vision   

Vision is one of the most important senses for living creatures for two main reasons:

• Perception

There is a need for them to identify objects and their location within their surroundings and so understand the environment in which they find themselves. This may be a very special ability as with insects, or an extremely sophisticated capability as with humans.

• Guidance

Sighted creatures have the ability to avoid collision with - or capture by - objects in their surroundings. They may also plan movements in their surroundings. For many too, there is a need to be able to position particular body parts in relation to each other and to locations within the surroundings. For example, with humans to see how their hands are positioned to pick up an object.

This suggests there may be an advantage to examining biological vision from two viewpoints:

• Spatial vision

The process for deriving information on spatial relationships.

• Temporal vision

The process for deriving information on changes within an environment over time.

In general, references to vision mean studies of the human visual response. Specifically, vision refers to the study of the physical processes involved in that response and so is focused on the eye. Visual perception refers to the study of the functional behavior of the physical and cognitive responses and studies the eye as well as the visual cortex within the brain. As a system, the human visual response accepts a stereoscopic image in general and as an output provides information to guide the behavior of a human.

There are many reasons why visual response should be studied. One is so that machine vision systems can be easily created; that is to say, a system that mimics all or part of the human visual response, or for that matter the visual response of another creature. With this base, though, response systems can be created with expanded capabilities. In particular, perceptive functions not part of human or animal vision. If that is the case, then any definition of images and image processing needs to be sufficiently broad to encompass all such possibilities.

There are many aspects of visual response that are of importance to psychology but rarely faced in the digital image processing literature. For example, when is an image perceived as being ‘natural’ and when as synthesized? Much more complex; how to measure the aesthetic qualities of an image? This is in fact a question of some significance in digital image processing as it would suggest means of improving the subjective quality. However, it is such a complex issue that in general the approach is simply to provide tools to a human operator and leave the issue to their judgment.

3. What Is An Image?   

The traditional view of an image derives heavily from experience in photography, television and the like. In the abstract, this view sees an image in these terms:

• It is a two dimensional structure.

• It is a representation.

• It is a structure with meaning to a visual response system.

While there are many implications of these terms, at this juncture it is useful to highlight just two. This definition assumes an image exists, but that in turn assumes a process of image formation. Such a process may influence an understanding of what is an image. There is also an implication in this definition of purpose. That is to say, the image exists so that it may be subject to some form of interpretative action.

This view of an image only accepts spatial variation. In modern parlance, this would be described as a static image. However, there are a number of potential applications where other forms of variation need to be considered. In accommodating those, the concept of an image needs to shift a little to encompass a view of a sample or section - a window - of a more complex structure. In particular:

• A dynamic image has spatial and temporal variation. In most contexts, this is usually referred to as video. In most cases this more complex structure needs to be viewed - if digital - as a sequence of images each representing a particular instance in time.

• Consider a volume. Then an image can be formed by taking a sampling plane through that volume and so the variation in three dimensions observed. This may be referred to as a volume image.

An image linked to a volume that changes with time is a further possibility. This has significance in medical image processing applications and some other areas.

4. The Relationship between Digital and Analog Images   

It is easy to become confused about the meaning of analog and digital with respect to images as there are many examples that seem to be neither one nor the other. A transducer is a device that converts some form of energy to an electrical signal. Here, the magnitude of the output directly relates to the strength of the input and the term analog was adopted to describe this. In time, though, analog came to mean any continuous signal. In a modern context, it is also used in the sense of being not digital.

To transform an analog into a digital image involves three conceptual stages:

• Sampling

A two dimensional grid is formed over the continuous analog image. Then the radiometric value is measured at the intersection points of this grid.

• Quantization

Each of the samples is restricted so that it can only take a finite set of values.

• Coding

These finite values or pixels (for picture elements) are now expressed by a binary number. Thus a digital signal is formed as a sequence of binary numbers.

The process is illustrated with a one dimensional example:

Figure 1a: A one dimensional analog signal

Figure 1b: The sampled analog signal

Figure 1c: The quantized sampled signal

The coding follows from the quantized levels.

Although the output is regarded as numbers, they would normally be communicated by a train of pulses. However, whereas in analog systems the shape of the pulse would be important and needs to be preserved, in a digital signal the issue is only whether what is present should be classed as a pulse or not. This illustrates a key reason for digitizing; it gives a result impervious to the usual sources of electrical noise so allowing perfect reproduction of an original.

Nyquist’s sampling theorem is a famous result that states a digital signal is equivalent to the analog from which it was derived if the sampling rate is at least twice the bandwidth of the analog signal. Equivalence means it is possible to perfectly recover the analog signal from the digital by a process of ideal low pass filtering. Hence in practice it is usually considered sampling needs to be 20% higher than the Nyquist rate so that recovery is possible with a practical low pass filter. If sampling occurs below the Nyquist rate, then any recovered analog signal will have an interference pattern termed aliasing. For this reason, it is important to filter an analog signal so that when digitized it is known to meet Nyquist’s criterion.

For images, aliasing can occur if the spatial sampling rate is not twice the spatial frequency response. That aliasing can take several forms. A very well-known one is ‘jaggies’ where straight lines are represented by closely linked segments.

Figure 2: An illustration of aliasing; ‘jaggies’ on a line.

This is a particular difficulty in computer graphics where regular geometric objects feature. In other images, a particular problem caused by aliasing is that it may lead to artifacts that can be interpreted as features of the image. This is a particular danger, of course, in medical image processing. Again, this problem can be overcome by filtering, in this case spatial filtering, the image prior to sampling. In visual terms, this is equivalent to blurring.

A discrete signal is not digital. The difference between the two is caused by the quantization process and it is an irreversible impairment. The impact of this is that the digital signal may be regarded as a noisy version of the discrete and so the original analog signal. The signal to quantization noise ratio is a widely quoted figure of merit and it is given by:

Signal to Noise ratio = 6.02N + 10.79 dB

where N is the number of bits of the quantized word. Tests can establish the minimum acceptable signal to quantization noise for different signals. Quantization noise, though, is not random and for that reason care needs to be exercised when examining any signal to quantization noise ratio. For common images, if the signal to quantization ratio exceeds 50 dB, then the human eye can rarely detect any perturbation which suggests that quantizing luminance to 8 bits is more than adequate. However, this is not true of all forms of images. For X ray images, for example, the far greater contrast ratio provided generally argues for quantization to 16 bits.

If the imagery is video there is a further issue to consider, namely aliasing effects due to temporal quantization. This leads to motion irregularities. An illustration can be seen in old western films during the inevitable chase. As the wagons increase in speed, the wheels appear to rotate faster, but then a point is reached where they slow, stop and then rotate in the other direction. This quantization effect is due to the video being a sequence of images and so, in the time domain, being discrete.

A static digital image, then, is a quite complex data structure with a range of geometric and radiometric attributes:

• Geometric

There is a finite spatial extent in each dimension defined by the size of the sampling grid in the image space. That is to say, the image is N pixels wide by M pixels high. This is sometimes termed the image resolution.

There is a spatial resolution in each dimension. That is, the spacing between sampling points in the grid. This can be uniform or non-uniform, but the former is more common. To illustrate, many common printers have a spatial resolution of 300 dots per inch (dpi).

There may also be an orientation of the grid with respect to some reference system that needs to be considered.

• Radiometric

There may be a single radiometric value, a triple of radiometric values representing red, green and blue primaries or many radiometric values representing a range of spectral bands. Each will have some radiometric resolution dependent on the quantization. In the case of a color triple, the representation may be RGB, but it can also be YHS (luminance, hue, saturation) or luminance with some particular set of chrominance coordinates.

There will be a contrast ratio set by whatever the smallest and largest pixel values represent for each radiometric value.

Depending, of course, on the type of the image, there may be some relationship – usually a power law – describing how pixel magnitudes map into brightness values. If it is a power law, then only the power needs to be known and this is described as the gamma.

5. The Concept of an Image Processing System   

An image processing system has four main sub-systems:

• Image capture sub-system

Some means of acquiring images. This may be as simple as a camera or as complex a system as a Magnetic Resonance Imaging ( MRI) scanner.

• Image storage sub-system

There is a general implication in (digital) image processing that images are stored both to allow processing and so they may be supplied at request to the destination response system.

• Image processing sub-system

A means of performing a range of processing operations.

• An Interface sub-system

A means of interfacing to a response system. This may be as simple as a display unit or a printer, or it may be, as for many machine vision systems, an interface to software performing some set of intelligent actions such as guiding a robot.

Special purpose image processing systems were manufactured in the past, especially for applications such as remote sensing. These were computers with a variety of hardware enhancements plus software that could take advantage of them. However, the performance limitations of computers which required that approach have largely disappeared, thus now an image processing ‘system’ is usually a software package with in some cases a few specialized peripherals. Some specialized systems are still manufactured though for machine vision applications.

6. The Process of Image Formation   

The traditional view of imaging centers on a lens of some form focusing energy onto a sheet located in the focal plane made of material sensitive to that energy- for example, in photography where the lens focuses visible light onto photographic paper. This view extends to electronic systems such as television where the outcome is dynamic imagery. However, there are many other models of image formation.

To illustrate, a scanning energy source and a sensor to detect the scattered radiation eliminate the need for a lens. Such an imaging system is important in a number of robotic applications where in this case the energy source is usually a laser. A slightly more complex version of this is offered by radar and sonar systems. Conceptually similar, but practically different, they have two possible modes of operation:

• Pulses are generated at some rate while the transmitter periodically scans a field of view. The reflected pulses provide information that can be used to generate a variety of images.

• A continuous signal is emitted. Due to the Doppler effect, objects moving in the field of view will cause the echo signal frequency to be changed where the size of the change relates to speed. This frequency shift is measured and used to generate images. Such Doppler systems are important in a number of areas such as weather forecasting and air traffic control.

These scanning systems may employ either a planar scanning beam or gather information over three dimensions.

A much more complex process of image formation is tomography. This has a number of forms such as CAT (Computer aided Tomography) and PET (Positron Emission Tomography) in medicine, and seismic tomography. The essence of tomography is illustrated in figure 3.

Figure 3: The basis of tomography

Here, an object – taken to be square for convenience – is to be imaged through some plane. That plane may be regarded as an array of elements. Some source may be used to direct energy beams through that object where the thickness of the beams will determine the resolution of the final image. Now if the magnitude of the source energy is known and a measurement is made on the other side of the object, then the attenuation is a function of the entire column through which that beam passes. If a series of such measurements are made over the grid, both rows and columns, then by applying mathematical techniques to the full set of measurements or projections, it becomes possible to identify properties of the individual grid points in the plane.

Magnetic resonance imaging offers a further variation. Here, the object of interest is placed in an intense magnetic field. That forces the molecules to align to the field. When the field is reduced, the molecules return to their normal state issuing radiation that may be detected. Thus the magnetic field can be designed to sweep the object with a two-dimensional wavefront and so an image is acquired again by sensors over a 360o field. Again, sophisticated mathematical techniques similar to tomography are used to derive an image. One of the attractions of MRI is that it is very well suited to forming a volume image and that is important in many medical applications.

7. The Image as a Representation   

In image processing, a scene means a three-dimensional arrangement of objects within a space illuminated by some energy source or set of sources. Thus an image can be a two-dimensional representation of a scene based on the reflectances of radiation off the objects within that scene. This describes, for example, photographs and television images. In this representation, depth and perspective are lost, but this presents only a minor problem to the human visual response. If depth information is a necessity, then it may be achieved through a stereoscopic system.

The image as a representation of a scene is extremely common, but it is by no means the only possibility. To illustrate, consider the three examples given earlier:

• For CAT, the image represents the permeability of soft tissue by X rays.

• For MRI scans, the image represents the density of soft tissue.

• For seismic tomography, the image relates to geological structure.

8. The Image Processing Hierarchy   

An image is just an array of data. Simple mathematical or logical operators may be applied to that array, but this may not achieve much of practical value. Therefore, more complex functions can be described built upon these simple operators. Then even more complex functions can be built on these. Thus image processing is best described by a hierarchy of four levels:

• Pre-processing

Strictly, this really should not be part of the hierarchy as its purpose is simply to correct problems resulting from sensing

• Low level processing

This level defines relatively simple mathematical or logical operators that perform basic actions on an image such as changing its contrast.

• Medium level processing

This level defines functions using sequences of low level operators to perform more sophisticated processing such as recognizing objects.

• Image interpretation

This level defines systems using sequences of medium level functions and perhaps low level operators that are intended to mimic visual response systems.

9. The Pre-processing Level    

Pre-processing revolves around models of image formation and is largely devoted to correcting problems that arise during the imaging process. Its principle concerns are these:

• Noise removal

Many sensors introduce random noise into captured data. The character of that noise varies according to its source, but in most instances it is additive white noise. Scratches in scanned film result in multiplicative noise meaning the effect varies with brightness level, and this can occur in other situations. Noise may also be colored meaning the contribution varies across the spectrum.

Noise may be reduced by filters. If the noise energy is low or the noise spectrum significantly different to that if the image, then a simple linear filter will suffice. If there is some concern about preserving edges in the image and the noise energy is a moderate, then a nonlinear filter such as a median or rank filter may be used. If the noise energy is a serious concern, then a more sophisticated approach such as Kalman filtering is required.

• Radiometric correction

Contrast, brightness or gamma may be corrected. This may require some calibration process for the imaging system. Figure 4 illustrates these actions. This image is of Lena, one of the most famous images in image processing and one of those almost always used to demonstrate advances in the field.

Figure 4a: The original image of Lena

Figure 4b: A change in contrast

Figure 4c: A change in brightness

Figure 4d: A change to a gamma of 2.2

• Registration

Registration is essentially limited to the various forms of remote sensing. Satellites and aircraft scan the earth’s surface, but they are subject to pitch, roll and yaw. This geometrically distorts the imagery. By linking the known position of reference ground points to their position within the imagery, it becomes possible to determine a correction function to overcome this distortion. That in turn allows accurate navigation through the data.

• Blur correction

There are two main forms of blurring that may occur in imaging. One is due to the lens employed. The other occurs when an image capture system has to function over a finite time. Then objects in motion may move a significant distance over that time.

Figure 5: Blurring caused by horizontal motion

Blurring is in effect a low pass filtering operation, but its correction requires some moderately sophisticated processing.

• Geometric transformation

In some cases, it is desirable to perform some form of complex nonlinear transformation on the image data. For example, to transform satellite remotely sensed data to a standard map projection. Such transformations are given several names although warping seems slightly favored. However, more recently warping has been used to refer to a process whereby one object is transformed in a video sequence to another; for example, as in some television commercials where a cat is transformed to a tiger or a man to a woman.

• Mosaicing

In satellite remote sensing and some other applications, an imaging system has a limited field of view but what needs to be processed is an image over a much larger field of view. Thus several images need to be joined together or mosaiced to form the final desired image. This can be a very complex procedure indeed as there can be overlap between the base images, and a complex range of distortions that need to be accurately removed before the connection is made.

• Color correction

A range of color distortions may arise in imagery. For example, color casts across the image may occur, there may be a need to transform the primary set into another or there may be a need to change the effective color temperature of the illumination.

• Re-sampling

There are many situations where the resolution of images needs to be changed. For example, when a 300 dots per inch image must be sent to a 150 dpi printer. This requires a process of re-sampling. It revolves around the two processes of interpolation and decimation.

Conceptually, the sampling rate of a digital signal may be altered by:

* converting the signal to analog;

* low -pass filtering if necessary;

* sampling at the new rate.

However, this may be done digitally as follows:

* Express the ratio of the two sampling rates as an integer fraction. Consider figure 6. The upper part of the figure shows the current set of samples of some one-dimensional analog signal and the lower the desired set of samples of the same signal.

Figure 6: A signal to be re-sampled showing the current and desired sampling points.

Thus in this case every four samples are to be reduced to three.

* Insert zero samples into the sequence according to this ratio. In this case, add a zero sample after each existing sample. That means there are now seven samples to be reduced to three.

Figure 7: The signal to be decimated with zero samples inserted and their location with respect to the desired sampling

* Now digital low pass filter - interpolate - so these samples assume values.

Figure 8: The interpolated signal.

* Omit samples according to the ratio to achieve the required result. In this case, for the sample taken, the next two are omitted and the process repeats. This is decimation.

Figure 9: The decimated signal

There are a number of specialized methods that may be used to reduce the computational effort in this process.

10. Low Level Image Processing   

The differences between low level processing and pre-processing are largely conceptual. Pre-processing may be described as image restoration with the objective of producing a result independent of the imaging process. It is almost always objective rather than subjective processing. Low-level image processing uses almost exactly the same techniques, but the objective is usually different. In general, its aim is to improve image quality thus it may be described in general terms as a process of enhancement. Given the similarities with pre-processing, the two stages are often merged when supply is from a simple image capture system.

The essential processing actions of this stage are again mathematical or logical operators directed at radiometric or geometric transformations of some kind. The common actions are based on the following:

• Spatial filtering

This can take many forms. However, a very common is edge enhancement as that can improve an image’s subjective qualities. It may also be used to remove or reduce the significance of objects within the image that are not relevant to the interpretative process planned.

• Histogram processing

A histogram is a count of the number of pixels in the image at each particular radiometric value.

Figure 10: A histogram of Figure 4a.

Histograms are extremely useful for a number of actions. For example, features can be identified and separated out. A histogram also states something on the contrast within the image. Thus histogram equalization can bring features out of shadow regions or otherwise improve the visual quality of the image. Figure 11 shows a very simple example of the value of histograms where all features outside a range have been set to black and those within set to white. The result is similar to, but not the same as edge detection.

Figure 11: Lena subject to two threshold limits.

11. Medium Level Image Processing   

Medium level image processing goes under a variety of names, but one of the more common is image analysis. In the abstract, the stage may be described as the sequence:

* feature extraction;

* pattern recognition.

Feature extraction has a considerable breadth of techniques. Much of it begins, though, with a segmentation process utilizing low level operators. This process is usually one of:

* level-based;

* edge-based;

* texture-based.

As this suggests, segmentation is based on some similarity principle applied to pixels of an image and its outcome is sets of such similar pixels. This principle can relate to any combination of spatial, spectral or temporal elements. The sets may sum to be the image or they may not, and they may be overlapping or not.

These sets may be the feature. However, more commonly a feature is a mathematical measure of some form derived from a set such as a count of the number of similar pixels. A more complex example is that of simple parameters derived from edges or contours such as:

* morphological measures such as moments or enclosing rectangles;

* topological measures such as the ratio of the area divided by the square of the perimeter.

Even more complex measures might be, for example, Fourier coefficients derived from the shape of the contour or edge.

Feature extraction may need to involve far more sophisticated operations. There are many applications, particularly with scene analysis, where the object of processing is to identify if particular objects exist within the scene. That suggests a process of correlation be applied. However, as three dimensional objects their exact appearance within an image will depend on camera position, distance to the camera, lighting conditions and so forth. That makes recognition very complex.

An ideal solution would be some transformation that is applied so that the outcome is invariant of such elements. No one transformation offers invariance to all of these, but a number offer invariance to particular elements. One of some significance in machine vision, for example, is the Fourier-Mellin transformation. It offers RST (Rotation, Scale and Translation) invariance, thus after transformation these influences do not exist so allowing a simple process of recognition of objects within a scene. A range of other transformations may be used to, for example, remove shading effects.

Pattern recognition is of two broad forms:

• Statistical

The recognition is to which class the object belongs (of some set of the defined classes) using the methods of statistical estimation.

• Syntactic

The recognition is a description of some form. The syntactic description is usually a string of symbols derived from the feature in some way. For example, a contour may be modeled as a connected set of primitive contour elements and then by a string where each symbols represents a primitive. Then recognition is achieved by identifying if that string belongs to a particular formal grammar using methods derived from automata and formal language theory.

Statistical methods are the more widely used for a number of reasons, not the least being there are some very sophisticated techniques available for estimation. In most cases, though, the estimation uses relatively simple and well-known techniques. For example, assume that a feature in an image is described by N numerical parameters. Then a given measurement of a feature in a particular image forms a point in an N-dimensional recognition space. If these measurements are made over a series of images, or features within images, then it would be expected there will be clustering in that N-dimensional space. Therefore, the space may be divided into zones, and from these a simple estimator formed for any later recognition. Figure 12 illustrates for a simple recognition process involving just two measured parameters.

Figure 12: A simple two parameter feature estimation process

This is just a simple form of supervised learning. When the parameters relate to natural phenomena such as in remotely sensed images where the interest is in, for example, crop growth, then the clustering can be quite poor. In this case, a simple set of planes in the recognition space will be quite inadequate and quadratic methods may need to be applied. If the clustering is very broad, then a nonlinear structure such as an artificial neural network may be needed.

This level of image processing is now very well developed with an extensive literature and repertoire of techniques. In many respects, it is the core of current image processing in almost all application areas.

12. Image Interpretation   

In an abstract sense, image interpretation is concerned with identifying the nature of the objects found in medium level image processing and then examining the relationships that exist between them. However, there are many differentiations that could be applied to this and the fact there is just the one term hints at the complexities of the subject. Many would also prefer a tightening of the definition in order to express this level as seeking an ‘understanding’ of an image. In general terms, this is concerned with scenes and so achieving understanding means gaining abilities equivalent to those of the human visual response.

Achieving a system comparable in performance to the human visual response is extremely difficult. Why that should be the case needs to be examined. Consider a photograph showing a woman with two young girls sitting on the grass in front of a building. Then:

• To an architect viewing this image, the information of interest is the building – its age, style, construction method and material, placement of windows, height, etc – and the rest is largely noise.

• To a clothing designer, the interest is the styles of the adult and child clothing, the materials, accessories and so forth.

• To a landscape architect, the information is the lawn on which they sit, the flowers around, and so on.

• To a health worker, the information is the apparent health of the woman and the children.

• To a historian, the information is who they are and their relationships.

An image as a data structure is relatively modest in a modern context - it is only about a megabyte or so - but as an information structure as this illustration shows, it can be enormously rich. Because it may be interpreted in so many ways, clearly image understanding must be within some context. That context both filters and prioritizes.

Interpretation of any form is based on evidence. Each piece is carefully considered in the light of what it may contribute to posing or confirming a hypothesis. Consider the previous example. The image outlined contains a representation of a human being. The size in relation to other objects in the scene plus the silhouette provide evidence on whether this human is a child or adult, and then whether that person is male or female. Of course, the clothes worn are usually an excellent indicator of gender and the details of the clothes will give further evidence of age. The bearing of the person, the texture of their face, the positioning of their arms and hands provide additional clues on age and relationships to the other figures in the scene. A second issue of interest might be is the building in the scene facing east or some other direction? Noting the shadows in the image and the general shading across the scene will give information on that. Further, if the sun was behind the photographer, then the presence of particular shadows will suggest what else might have been at the rear.

This highlights a very critical point. Image interpretation has to work at various scales in the image. Shadow effects showing the positioning of the building are gradual and occur only at a coarse scale. Folds of clothing showing the orientation of a person are determined at a much finer scale and the determination of texture of their clothes at an even finer scale. In similar fashion, a tree is evident at a coarse scale, but it is the details of the leaf and branch structure that often determine what type of tree it may be. Thus the focus of image interpretation must be on identifying features gathered at various scales and then reconciling them.

A complicating factor is that scenes in particular are usually ambiguous. That is to say, there are multiple possible interpretations of sets of features identified within the image. One reason for this is occlusion; part of an object is hidden by another, a very common characteristic of most scenes. Thus an interpretation of an image of a scene depends on a set of constraints being defined to guide the interpretation. In general terms, it is extremely difficult to find some set of constraints that can achieve much more than reduce the set of possible interpretations to a small number. For that reason, the constraints need to be augmented with a set of heuristics whose task it is to assign probabilities to each interpretation and in that way derive the most likely. That is then the outcome of the image understanding process.

Given this, high level image processing:

• can be seen as a process of ambiguity reduction;

• since it does involve constraints, it involves non-image data and that data is often far more important to the success and objectives of the overall task;

• since it involves heuristics, requires a very good understanding of the real world environment from which that image was derived.

Clearly, at this level image processing is very closely linked to artificial intelligence.

Serious research into machine-based image understanding began over quarter of a century ago. The problem of understanding line drawings was examined as they were easily produced and seen as a likely group of restrictive images from which general conclusions could be drawn. Further, the task had considerable industrial importance. This leads to an interest in interpreting images of regular structures where the interpretation was actually of the edges within such images. Much was done in this area as well, but it was found very difficult to translate results to other images.

A more formal definition of image understanding shows that much of it may be described by an NP-complete problem known as the consistent labeling problem or constraint satisfaction problem. The computational effort needed to gain an understanding is therefore combinatorial and that makes it an exceedingly complex problem to solve.

13. Interpolation in Image Processing   

Although low level suggests simplicity, in fact a number of low level image processing functions are extremely sophisticated in their own right. Interpolation, at the heart of much geometric processing, is a good example of this. It will be illustrated by discussing the actions involved in rotating an image.

A viewer using an image processing package usually observes test images on a rectangular display device. Each displayed image derives from an array of numbers stored in a computer’s memory. When that viewer requests rotation of an image, what they expect to see is the displayed image subject to the rotation they have requested. However, the displayed result still derives from an array of numbers as before. Clearly, a mathematical operation must be performed on the array of numbers forming the original image such that it is transformed to a new array that exhibits the rotation.

Conceptually, in rotating an image the following occurs:

• The original image is defined on some sampling grid that for convenience may be taken as the display grid.

• That image is rotated by the angle required. As can be seen from Figure 13, the sampling grid of the rotated image is such that sampling points do not correspond in general with those of the display grid. Indeed, it will be noted the position of each rotated sampling point to the nearest points in the display grid is quite variable.

Figure 13: The relationship between the sampling grids of an image and that image rotated.

• The mathematical operation required to compute each pixel of the display grid from the rotated image may be described in this way:

* A smooth surface needs to be constructed for the rotated image.

* That surface is sampled at the points of the display grid to produce the desired result.

That is to say, the operation required is interpolation.

As this description suggests, rotation can be computationally a very expensive operation. In the early days of image processing, the interpolation was often simplified in order to reduce in particular the number of multiplications required. As a result, interpolation may be said to have progressed through three eras in image processing:

• Consider the shaded area in Figure 13. Each pixel to be computed falls within a cell of four pixels of the rotated grid. Then the simplest approach is to identify the nearest pixel within the cell to the desired sampling point and use that as the output value. This requires very little processing, but the resulting image can be severely distorted. Figure 14 shows an image rotated by 30o on this basis. If it is examined carefully, then distortions can be seen along edges and in regions of texture. Nevertheless, if the aim is just subjective, the result can be acceptable.

Figure 14: Lena rotated through 30o with nearest neighbor interpolation

• A slightly better approach - and certainly a vast improvement aesthetically - is to fit a plane to the four pixels in the cell and sample that. Figure 15 illustrates.

Figure 15: Lena rotated through 30o with simple neighborhood interpolation

• If the intention is to use the rotated image for objective purposes, then the best approach would be to interpolate using a two-dimensional digital filter as close to an ideal low pass filter as possible. However, that is too complex. For that reason, by far the most common approach at present is to use an interpolating function that leads to as smooth a surface as possible while involving the smallest neighborhood. That is to say, to use spline interpolation. Splines are very well understood as they play an extremely important role in computer graphics.

14. The Edge Detection Problem   

A uniform image has no information. It is the change within an image that provides information and thus the means of detecting that change are central to most medium and higher level image processing functions. As an aside, the human eye easily interprets images consisting solely of change information, namely drawings or sketches.

Consider this definition:

An edge is a boundary that represents an area of change or transition within an image.

It is attractive in that it is inherently two dimensional and it assumes an edge is a structure with area and other properties. It also highlights why there is such confusion over edge detection as this definition may be used in many contexts. These may be divided into two broad divisions concerning in a wider sense what change or an edge represents. Two examples illustrate the issues- first, a typical industrial problem where parts are machined and then inspected by an electronic vision system. In this case, lighting can be carefully controlled so that the significant change within the image corresponds to the physical boundaries of the viewed parts. Thus a simple process of differentiation and thresholding identifies all edges of importance for the application. That information may be used to identify if the work that should have been done has been. In this instance the detection of edges of interest only requires low level image processing activity.

Contrast this to scene interpretation. A scene is a representation of a number of objects located in a three dimensional space. That space will probably have several lighting sources; thus there will be a complex shadow pattern. Examining that – which means examining the scene at a coarse level – will provide information on the orientation of objects within the scene to one another. Each of the objects is likely to be highly textured, in part due to lighting and in part due to the textures of the surfaces involved. To locate this information requires examining the object to gain clues on its nature and shape, and then examining each of the surfaces to identify what they may be. Thus change has to be detected at multiple scales in such an image and then interpreted; hence the edge detection problem shifts to being a medium or high level image processing activity. Naturally, it will use the simple operators of the low level machine vision task, but there is a need for additional functionality. Overall, then, this variation in the outcomes desired or needed makes edge detection a very good illustration of the distinctions in the image processing hierarchy and how medium and high level functions are constructed.

A first issue to raise with edge detection is what should be the outcome of a detection process? This is not as obvious as it seems. Recall that a digital image is defined on a sampling grid. Figure 2 illustrates a situation where an edge in the original analog image falls between the sampling points of the digital. For that reason, the outcome of an edge detection process can be of two forms. It can be a geometrical entity independent of the sampling structure that defines the edge-that is to say, the straight line shown in Figure 2. On the other hand, it may be some marking or classification of image points to indicate they are adjacent to or part of an edge- that is to say, the segmented line shown in Figure 2. The second is more favored. However, note an important consequence. The edge described in such a way is likely to a representation that may be described as the real edge possibly suffering significant spatial quantization noise. That in turn may disturb later interpretative processes.

Most people questioned on what is an edge would probably define it as a discontinuity in luminance. Accepting that for the moment, consider some one dimensional examples of such discontinuity shown in Figure 16:

Figure 16: Some examples of simple discontinuities

The first of these is an ideal step edge, the second an ideal ramp, the third a valley and the fourth a peak. These discontinuities differ in their mathematical properties. For a step, the discontinuity is in the luminance function itself and so derivatives are not defined at the edge point. For the others, however, there is no discontinuity in luminance, just a change, but there is in the first derivative; thus they are only variations on each other. Thus if discontinuity is taken as the criterion, it has to be concluded that the concept of edge may not be described by a single property, but rather that a pixel is classified as an edge point if it exhibits one of a set of properties. That in turn suggests that there may be a need for a range of (low level) edge operators oriented at detecting particular types of edge points in an image.

Now consider the two structures shown in Figure 17.

Figure 17: Two simple, ideal edge structures

Both of these are often seen as simple, but ideal forms of edges and so a single structure. Indeed, they are sometimes referred to as line edges. However, both have two discontinuities, hence the question is: should these be identified as an edge or not? Further, there is the question of how should they be identified? That is to say, should an operator be created to identify them at a low level of image processing, or should the simple discontinuities be detected and an interpretation made of this information at a higher level? An issue in both cases is how large should the ramp be or how wide the level before they are seen as image features in their own right and not edges?

These examples are extreme idealizations. Real edges are much more complex. Figure 19 shows a profile taken across the image of Figure 18 along the line shown.

Figure 18: An image with a typical edge structure

Figure 19: A profile taken on the line shown in Figure 18.

Clearly, there is considerable change across this profile. However, note that there is a structure here of larger scale changes blended with very localized change. Therefore, if a low level operator is applied that simply detects change, quite a string of pixels will be identified as putative edge points about what a viewer would see as ‘the edge’. This emphasizes two points- first, again the importance of scale in examining the edges of a complex image. Second, that while edge detection may be seen as a simple task, in many cases it requires quite a complex process of interpretation.

The definition given earlier for an edge suggests it is important to explore two properties of edges before formulating any edge detection system. An edge is a geometrical structure formed usually from some complex set of curves. It also has spatial properties in the form of its width at each point in the structure as well as a magnitude over that width where this relates to the degree of change. Just to complicate the issue, it may be noted there are circumstances where edges form an implied geometry and so have null spatial properties. For example, consider the well-known visual illusion shown in Figure 20.

Figure 20: A triangle visual allusion.

For some image understanding applications at least, it can be important that the implied structure is recognized as certainly the human visual response reacts to it.

The spatial properties of an edge define what type of edge it may be, for example, some primitive description such as step or ramp. However, note the emphasis of the earlier definition. It is focusing on the edge as a structure having the overall property of being a transition. That definition encompasses the prospect that over the entire edge these spatial properties can change. This is often a problem with traditional low level edge operators. They only accept particular conditions to identify edges and that means part of the higher level of interpretation is joining together detected edge segments. Thus it is better in the main to focus in edge detection on preserving geometry – or more truthfully connectivity - and separately identify the form of the spatial characteristics.

The geometry of edges is surprisingly simple. Edges may be seen as complex combinations of three primitive forms namely curves, intersections and polyjunctions. These are illustrated in Figure 21.

Figure 21: Primitive edge geometrical elements; the curve, intersection and polyjunction.

Now consider the structure of Figure 22.

Figure 22: A simple edge geometry of two curves and one inflection point.

In many higher level applications, the actual geometrical properties of the curves are not very important, but the position and form of the inflection or junction point as shown in Figure 22 is. Indeed, studies dating back to the 1950s show inflection points have very profound perceptual significance and are a key element in human visual recognition. For that reason, an alternative way of describing edges is by the nature of the inflection points within them. And aside; many current edge detection operators have great difficulty with inflection points. They may distort their geometry or even break the curve. For that reason, much work has been done on systems to just detect these points.

When edge detection needs to be a higher level image processing function, then it tends to be a process that includes all or most of the following actions:

• Pre-processing usually directed at reducing noise.

• Scaling to separate the target image into a set of scaled representations.

• The identification of putative edge points in each of the scaled representation by edge detection operator. That is to say, a low level image processing function is applied.

• There may be a need for localization to correct for distortions introduced in detection, particularly at inflection points;

• Often thinning need to be applied to reduce edge structures to lines.

• Linking may need to be performed to identify incorrectly classified points and also join edge segments.

• Finally, combination is required to join the scaled estimates and produce a consolidated edge map.

Some of these are still the object of research and so their exact form is open to question. Note that scaling is usually by a Gaussian filter as it has some very desirable properties. Further, once the consolidated edge map is produced, then, depending on the global objective, it would normally be passed to an interpretive system.

One of the areas of this system to emphasize is linking. Much edge detection is point-oriented, but an edge is a structure. Therefore, point processes are likely to miss true edge points due to particular inadequacies. Linking seeks to rectify that. While this sounds a simple task, in general it is not.

Edge detection operators are usually of three types:

Edge strength

Orientation

Model fitting

An edge strength detector usually employs either a gradient or directional derivative operator. However, in recent years some success has been reported with statistical methods based on covariance. Whereas edge strength methods are focused on determining if there is change and where it is a maximum, orientation methods focus on determining where there is minimum change. Typically, they are based on a line segment and how to orient it via correlation or similar methods so that it fits along the edge. Model fitting methods assume that an image may be described by a generational model. This may be a statistical model or a surface model. In the latter case, a surface is fitted to available data; then it is differentiated to form a gradient. In that respect, it is quite similar to edge strength methods.

Of these methods, edge strength tends to dominate. Figure 23 shows a typical example.

Figure 23: Edge detection through a simple edge strength detector.

Early and relatively simple methods continue to be employed such as the Sobel or Roberts operators. However, more sophisticated medium level functions based on the Marr-Hildrith approach, and in particular, Canny’s approach, are the choice for more complex applications. The development of edge detection operators continues to be a topic of research interest leading to, for example, the interesting Iverson-Zucker system. It needs to be mentioned, too, that there are some useful methods based on relaxation techniques. While these offer promising results, there has long been a problem in that these are computationally expensive.

15. Applications of Image Processing   

As an enabling technology, image processing grows and strengthens through the many diverse applications that employ it. They are quite remarkable in their diversity and that makes this review of the applications a rather brief and limited overview:

• Machine Vision

Machine vision is the ‘industrial arm’ of image processing. Like many others, it is concerned with the processing of images derived from a scene. It differs, though, in that the scene is simple. Further, the lighting is usually controlled, occlusion avoided and other actions taken to greatly simplify the image processing tasks. Generally, but not always, that image processing begins with the detection of edges which, through the controlled lighting, identify the actual physical edges of objects.

While the obvious application area is robotics, in practice very few robotic systems need a visual sense. The more important area is inspection systems of various kinds. There are many possibilities for these. For example, an inspection system to check that all components exist in an assembly such as on an electronics printed circuit board or that boxes are properly packed. Alternatively, to ensure a machining system has produced a part to specifications.

A rapidly developing area is guidance systems which differ from the above in that the scene can be complex. Traditionally, these have been for manipulators of various kinds. The area that is growing most rapidly, though, is automated guidance systems for vehicles of all kinds.

• Image Compression

Image compression is not obviously image processing. The objectives of compression are to reduce the data volume of an image or video stream. That allows the data to be transmitted over a lower bandwidth channel as in digital television or for more information to be held on a storage system like DVD. It follows there is a de-compression or reconstruction stage after transmission or storage to recover the original data.

Compression can be error free so that an exact copy of the original can be re-created, or it may be lossy. Lossy means little subjective difference between the original and re-constructed image but much objective. The advantage of it is that the compression ratio is much greater. Lossy compression methods are described in standards such as JPEG and the various MPEG standards used for the transmission of images over the Internet, direct storage on DVDs and in digital television systems.

• Medical Image Processing

Medical image processing is a vast and rapidly developing field. The source of the imagery is usually from scanners delivering volume data. The purpose of processing is to provide information for diagnosis, to identify particular patient conditions and observe the impact of various treatments. However, there are evolving applications. To illustrate, recently there have been interesting experiments combining image processing, graphics and virtual reality. Here, the patient is scanned and using image processing techniques the patient’s organs identified. In the operating theatre, the surgeon uses virtual reality glasses to see the reconstructed organs within the patient’s body. Prior to operating, the surgical team may have identified procedures to follow, thus information relating to this may also be shown.

• Security applications

Image-based security systems can be quite simple. For example, observing an area divided into zones and setting off particular alarms when there are certain types of activity in each. However, the two areas currently at the leading edge in this area are human face recognition and handwriting analysis, and both are quite complex.

Face recognition seems to be relatively simple. The position of the eyes, nose and mouth are difficult to change, thus parameters based on ratios of measures of these would seem sufficient. However, a face is three dimensional and attempting to recognize one with accuracy from two-dimensional imagery that can be of relatively low definition is quite difficult. In spite of much research and development effort, existing systems cannot be described as performing to the level needed.

Handwriting analysis has a long history. The ability to interpret handwritten documents is clearly attractive in a number of applications, and vital to nations where the writing systems used by their citizens do not fit well with the alphanumeric keypads and similar devices of the West- that is to say, those who use pictographic writing forms such as China and Japan. As an image processing task, handwriting analysis is essentially an image interpretation task but much more focused.

• Entertainment

The entertainment industry has begun to rely quite heavily on special effects over the last decade. While these in the main have been achieved through graphics, image processing is also important. Much of this might be described as live animation. That is to say, the basic data derives from a real object, but the processing which follows owes more to graphical techniques than many traditional image processing methods. To illustrate, many television commercials now have animals or babies that smile, wink or perform other simple actions. A more complex activity is where scenes of a deceased actor are processed to extract their face and mouth in various positions so that a composition can be made of them acting entirely new scenes. Again, commercials pioneered much of this. As graphics develops more realistic ways of synthesizing humans, this application may fade.

• Image databases

Image and video databases have considerable potential, but have been exceedingly difficult to develop for two principal reasons. One is the lack of theoretical development at almost every level as they are very different to conventional databases. The other is the computational demands. A query in such a database may be on a simple property but it is more likely to be on a derived property. Thus in a general sense an image or video database is a content addressable database.

Various taxonomies have been suggested for possible image databases. One that is particularly relevant to this chapter relates to the image processing hierarchy given earlier and leads to a taxonomy with four elements.

An image may be processed to derive a set of indices. These can be, for example, the dominant color or a spatial or spectral histogram. Therefore, the lowest level of image database is one that simply searches on these indices. As they can be easily implemented with standard relational databases, there are a number of examples of them available now.

While such systems seem limited, an examination of a number of them shows they can be surprisingly effective. Their mode of operation also suggests some important characteristics for future systems. In general, they request a very simple attribute of the target image from the user. They will then present a range of images that meet the criterion, but which differ significantly in other respects. The user is asked to identify which best meets their criteria and a second, more precise search follows. In the abstract then, the principle is multiple searches where the initial are focused on filtering and the latter on refining a target set. Hence a potentially excessive computational burden is significantly reduced.

A more sophisticated approach needs to recognize an image in general terms has some set of elements together with a structure which shows how these elements are interconnected. Then the next level in the hierarchy is simply to identify if a nominated element exists in a given image under test. This is essentially pattern recognition. A further step up is to identify if a particular sub-structure exists where this is a distinct but nonetheless higher form of pattern recognition. The final level is based on interpreting the structure and that is image understanding.

• Remote sensing

While there are several forms of remote sensing, in a modern context it tends to mean satellite remote sensing with some extension to aircraft-based systems. It is an extremely important activity for monitoring the environment which means it is also a vital tool for the primary industries.

In general, multiple coincident images are captured in remote sensing, each covering a particular spectral band usually ranging from the visible to the infrared. The number of bands can range from five on some older satellites to over 2,000 on some proposed. Such bands give information related to chemical properties. Radar remote sensing systems in contrast provide information on physical structure where this varies according to the sensing frequency.

Satellite remote sensing began as a largely interactive activity. Over time, though, theoretical developments have seen much remotely sensed data being an input to computational models. For example, much meteorological information is now derived from climate models processing satellite data.

Remote sensing systems require very sophisticated pre-processing structures. These, for example, remove atmospheric absorption and scattering effects, plus geometric disturbances caused by the spacecraft’s pitch, roll and yaw. Since the satellite orbits vary, another very important but computationally expensive task is to map the data to a standard projection so that sequences of images taken days, months or years apart can be compared.

Much of the processing focuses on manipulating spectral bands to highlight vegetation growth, sea surface temperatures, algae levels in rivers, particular forms of mineralization and so forth. A very fruitful development, especially for mineral exploration, has been to combine radar and spectral images plus other data.

16. Some Image Processing Packages   

A very comprehensive source of information on packages – and indeed all aspects of image processing – is maintained on the Computer Vision homepage supported by Carnegie Mellon University in the U.S. Most image processing packages are targeted at specific applications such as machine vision applications or medical image processing. However, this web site lists are large number that are not ranging from packages intended for research applications to more simple tasks.

Probably the most important of these are usually described as toolkits and the web site lists over 50 that are available. Generally, these have some base software platform that provides format conversion and a number of very common processes plus some means to extend the functionality via modules or plug-ins. A selected set of examples of these are as follows:

• A much-favored toolkit is PBM (now PBMPlus). Developed for the Unix environment some years ago, it consists of a rich collection of functions that may be united via a script to perform complex image processing tasks.

• Khoros is a similar package, but it has been created in the form of a programming environment. Again, through that environment users can form complex image processing actions. This is achieved through a visual programming language where the user manipulates screen objects.

• Part of image processing is the manipulation of images for aesthetic or publishing reasons. Hence there are a number of commercial packages available such as Adobe’s Photoshop that do this very well. Photoshop allows third party plug-ins and a number have been developed for more general image processing such as Fovea 2. Those with a high level of programming skills and some experience may consider downloading the Photoshop SDK (Software Developer’s Kit) and develop their own.

• A freeware package on similar lines is GIMP (for GNU Image Manipulation Program.) It is Unix-based, but there are versions for other operating systems. As yet there are few general image processing plug-ins, but their development may be of interest to skilled programmers.

• On partly similar lines, an image processing module is available for the popular Matlab package. The module covers a range of common functions, but given the nature of Matlab, it is easy to develop others for specialized purposes.

A slightly different approach to the toolkit has been followed in the development of the Image package by the National Institutes of Health of the United States. Originally developed for processing biomedical images, it had to face the problem of the extreme diversity of such images. Rather than form a toolkit, it solved that by including a macro language similar to the Pascal programming language so that users could create their own medium level processing routines. This made it very popular indeed and it became widely used for applications outside of medical image processing.

More recently, the NIH developed JImage, a Java version of Image that will run on any machine that supports the Java language, and very few do not. Being written in java, it is very easy for users to add their own java routines and again extend the package to meet their particular needs. This package is an excellent tool for those wishing to experiment with image processing. Indeed, many of the images within this chapter were processed using it. It is available free of charge.

For enthusiastic programmers, Java has a set of image processing libraries from which an application can be constructed. These are, in the main, directed at graphics applications, but there are quite enough for serious image processing tasks. There are also quite a number of C image processing libraries available both free and commercially from which applications may be constructed. It may also be noted that simple image manipulation packages are becoming increasingly common with the rise of digital photography.

Related Chapters   

Click Here To View The Related Chapters

Glossary   

Analog image processing

 :

image processing using lasers and optical components.

Byte

 :

a standard data element in computing consisting of 8 bits

CAT

 :

Computer Aided Tomography or Computerized Axial Tomography

Digital image processing

 :

image processing via software and computers.

Image resolution

 :

the number of pixels vertically and horizontally in an image

k

 :

kilo as in computing meaning 1024, as in a kilobyte

M

 :

mega as in computing meaning 1024x1024 or 1,048,576, as in a megabyte

MRI

 :

Magnetic Resonance Imaging

Pixel

 :

picture element, a point in a digital image, also termed a Pel

PET

 :

Positron Emission Tomography

RST

 :

Rotation, scale and translation (invariant)

Spatial resolution

 :

the spatial extent in units such as centimeters of an image.

Bibliography   

A very large number of books are available on all aspects of image processing as a check with an on-line bookstore or search engine will show. A very select group of references indeed are the following:

Gonzalez, R.C., Woods, R.E. (2001)"Digital Image Processing." 2nd Ed. Prentice Hall.[A very popular introductory text covering most of the fundamental area].

Pratt, W.K. (2001) "Digital Image Processing." 3rd Ed. Wiley.[Another very popular but more advanced general work.]

Batchelor, B.G., Waltz, F. (2001) "Intelligent Machine Vision: Techniques, Implementations and Applications." Springer .[A more advanced work on machine vision.]

Bovik, A.C. (2000) "Handbook of Image and Video Processing." Academic Press.[For the more advanced reader, but an excellent starting point for further studies or locating detailed information]

Davis, L. (Ed). (2001) "Foundations of Image Understanding." Kluwer Academic Publishers .[A good reference for this advanced topic.]

Veltkamp, R.C., Burkhardt, H., Kriegel, H. (Eds) (2001) "State-of-the-art in Content-based Image and Video Retrieval." Kluwer Academic Publishers .[More for researchers, but a good reference for this particular application.]

Ablameyko, S., Pridmore, T (2000) "Machine Interpretation of Line Drawing Images: Technical Drawings, Maps, and Diagrams." Springer .[There are a number of older books on this subject, but there has been significant progress in recent years. This is a better choice for reviewing developments in this field.]

Dougherty, E.R., Astola, H.T. (Eds) (1999) "Nonlinear Filters for Image Processing." SPIE Optical Engineering Press .[A very good reference indeed for advanced filters, particularly for noise filtering.]

Richards, J. A., Jia, X. (1999) "Remote Sensing Digital Image Analysis: An Introduction." 3rd ed. Springer.[Like a number of books on remote sensing, this covers a wide range of topics in both low and medium level image processing.]

Duda, R.O, Hart, P.E., Stork, D.G. (2000) "Pattern Classification." 2nd ed. Wiley,.[A great favorite on this topic.]

Costa, L., Cesar, R. "Shape Analysis and Classification." CRC Press, 2000. [An important reference for machine vision studies.]

Schowengerdt, R. (1977) "Remote Sensing: Models and Methods for Image Processing." Morgan Kaufman.[A very good study on the more advanced aspects of remote sensing including valuable information on pattern recognition]

Schalkoff, R.J. (1989) "An Introduction to Digital Image Processing." Wiley.[A good introductory text with a strong emphasis towards machine vision.]

Sullivan, R.J. (2000) "Microwave Radar: Imaging and Advanced Concepts." Artech House.[A highly technical book on the current state of imaging radar systems of all kinds including ground penetrating and synthetic aperture.]

Some of the key journals in this field are as follows:

Image and Vision Computing, Elsevier

Medical Image Analysis, Elsevier

Pattern Recognition Letters. North Holland

Pattern Recognition, Elsevier

IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE

IEEE Transactions on Image Processing, IEEE

Computer Vision and Image understanding, Academic Press

Two web sites worth visiting are:

The computer vision home page found at: http://www-2.cs.cmu.edu/~cil/vision.html [This page has an extensive range of resources devoted to all aspects of image processing covering both researchers and practitioners.]

The University of Southern California’s Signal and Image Processing Institute maintains an archive of the standard images used in image processing publications. It is found at http://sipi.usc.edu/services.html

Biographical Sketch   

Dr. Myers graduated from the University of Western Australia with a B.E. degree in communications engineering in1969 and an M.Eng.Sc degree in control systems engineering in 1971. He then worked for the Weapons Assessment Unit of the Department of the Navy investigating various combat systems. He later joined the Western Australian Institute of Technology where in 1987 this became Curtin University of Technology. In 1982 he completed a Ph.D at the University of Western Australia on image processing, presenting a thesis on filtering techniques derived from observations of the human visual response. About that time, he was strongly involved in the development of one of the first NOAA satellite receiving stations in the world. This later resulted in the formation of WASTAC, a remote sensing consortium involving Curtin and government agencies with a charter to collect satellite data. WASTAC now runs receiving facilities for several satellites and plays a pivotal role in Australian remote sensing. A first application of this work in remote sensing was a major study of how satellite data could be used to assist the fishing industry. Later, he was principal investigator of a study into the recognition of Australian wheat varieties through image processing techniques. More recently, he assisted in the formation and development of multimedia group under the Australian Government’s Cooperative Multimedia Centre scheme. Dr Myers has published a book on digital signal processing and contributed chapters to books on image processing and remote sensing. He has published a number of papers on the applications of image processing, especially in remote sensing and agricultural applications.


To cite this chapter


©UNESCO-EOLSS Encyclopedia of Life Support Systems