WO2000022728A1 - Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output - Google Patents

Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output Download PDF

Info

Publication number
WO2000022728A1
WO2000022728A1 PCT/SG1998/000081 SG9800081W WO0022728A1 WO 2000022728 A1 WO2000022728 A1 WO 2000022728A1 SG 9800081 W SG9800081 W SG 9800081W WO 0022728 A1 WO0022728 A1 WO 0022728A1
Authority
WO
WIPO (PCT)
Prior art keywords
block
coefficient
elements
architecture
filter
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/SG1998/000081
Other languages
French (fr)
Inventor
Rakesh Malik
Puneet Goel
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
STMICROELECTRONICS Ltd
STMicroelectronics Asia Pacific Pte Ltd
STMicroelectronics Pte Ltd
STMicroelectronics lnc USA
Original Assignee
STMICROELECTRONICS Ltd
STMicroelectronics Asia Pacific Pte Ltd
STMicroelectronics Pte Ltd
SGS Thomson Microelectronics Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by STMICROELECTRONICS Ltd, STMicroelectronics Asia Pacific Pte Ltd, STMicroelectronics Pte Ltd, SGS Thomson Microelectronics Inc filed Critical STMICROELECTRONICS Ltd
Priority to PCT/SG1998/000081 priority Critical patent/WO2000022728A1/en
Priority to DE69821144T priority patent/DE69821144T2/en
Priority to EP98950601A priority patent/EP1119909B1/en
Publication of WO2000022728A1 publication Critical patent/WO2000022728A1/en
Anticipated expiration legal-status Critical
Priority to US10/968,822 priority patent/US20050193046A1/en
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03HIMPEDANCE NETWORKS, e.g. RESONANT CIRCUITS; RESONATORS
    • H03H17/00Networks using digital techniques
    • H03H17/02Frequency selective networks
    • H03H17/0223Computation saving measures; Accelerating measures
    • H03H17/0225Measures concerning the multipliers

Definitions

  • the invention relates to area efficient realization of coefficient block [A] or architecture [A] with hardware sharing techniques and optimizations applied to this block.
  • the block [A] is connected to coefficient lines CLin_0,
  • [F] to be connected to perform filtering operation or a mathematical computing operation with optimization in hardware and provides a zero latency output.
  • the invention also gives the area minimal realization of digital filters based on coefficient block[A], when operated in bit serial fashion.
  • the optimization techniques and structure of the present invention are good for linear digital filters typically a finite impulse response(F-- ) filter, infinite impulse response filter(IIR) and for other filters and applications based on combinational logic consisting of delay element(T), multiplier(M), adder(SA) and subtractor (SS).
  • FIG 1 shows the field of invention, applications of the device
  • FIG. 1 shows the symbol of components used in device.
  • FIG. 3 shows the description of the components used in device.
  • FIG. 4 shows the bit serial FIR filter implementations.
  • Figure 5 shows an example of FIR filter.
  • Figure 6 shows one of the known minimization technique due to symmetry .of coefficient.
  • Figure 7 shows the structure of existing/known implementation technique for example FIR filter.
  • Figure 8 shows the generalized structure of existing/known implementation technique of coefficient block.
  • Figure 9 shows a realization of the coefficient block for coefficient value close to power of two.
  • Figure 10 shows the optimization (a) for realizing the coefficient block for example FIR filter, of the present invention.
  • Figure 11 shows the optimization (b) for realizing the coefficient block for example FIR filter, of the present invention.
  • Figure 12 shows the optimization (c) for realizing the coefficient block for example FIR filter, of the present invention.
  • Figure 13 shows the generalized optimized structure, of the present invention.
  • a serial coefficient multiplier(M) can be implemented by shift register using [T] elements and adder element [SA] (One shift means multiply by factor of 2). As shown in “ Figure 3" of the drawings, the multiplier is formed by adding the outputs corresponding to ones in the binary representation of the coefficient. Delay (Z "1 )
  • Delay by one frame of data is done by shift register (series of Flip-flops (T) connected to store and shift the input frame).
  • the number of Unit delay (T) in one delay element is equal to the frame size of the input.
  • Figure 4" shows the existing structure of bit serial FLR filter with coefficient lines CLin_0, CLin_l, CLin_n and the coefficient block [A] having the coefficients c(0), c(l), c(2),....c(n).
  • the coefficient block is connected to delay element [Z '1 ] and serial adders [SA] to form filter structure.
  • Y(n) c(0) X(n) + c(l) X(n-1) + c(2) X(n-2) + c(n) X(0)
  • coefficient lines CLin_0, CLin_l, CLin_n are common and connected to input X[n].
  • the output lines CLout_0, CLout_l, CLout_n are connected to block [E], consisting of delay element [Z "1 ] and serial adders [SA] elements.
  • the structure makes easy realization of share-able multiplier in the coefficient block [A].
  • An example of share-able multiplier with coefficient values 3,11 is illustrated in " Figure 4". The realization of these coefficient separately would require 4[T], 3[SA] elements.
  • CLin_0, CLin_l,... being common, the hardware is realized using 3[T], 2[SA] elements.
  • Another feature of the structure is that the structure inherently requires more storage area, represented by [Z "1 ], as compared to implementation 2, since the storage is done after the multiplication.
  • the storage area of each delay element [Z "1 ] is (m+n).
  • the total storage space of the delay elements is (m+n) * (number of coefficients -1).
  • the coefficient line CLin_0, CLin_l are not common.
  • the realization of coefficient block [A] using share-able elements is not present.
  • Another feature of this structure is that it inherently requires lesser storage space, represented as [Z "1 ], as unlike in previous implementation, here the storage is done before multiplication.
  • the storage area of each delay element [Z '1 ] is (m).
  • the total storage space is (m) * (number of coefficients -1).
  • the invention reduces the hardware of the coefficient block [A] by having shareable elements in coefficients, even if the coefficient lines CLin_0, CLin_l, are not commonly connected.
  • the share-ability of hardware in block [A] is a limitation.
  • implementation 2 is area efficient with " respect to implementationl due to reduced delay elements size. Over and above this by having share-able multiplier or reduced coefficient block [A], which are the key features of the invention, implementation 2 becomes still more area-efficient. This reduction is extendable to other filter based on coefficient block [A] as stated in the first section.
  • the present invention operates on integer valued coefficient. Further, to quote Norsworthy and Crochiere (Delta-Sigma Data Converters LEEE press pp-435, copyright 1997)
  • Bit serial architecture reduce the interprocessor communication down to 1 bit.
  • the number of processors is very large, but because each processor is so small, the overall economy is very high.
  • Bit serial architectures are usually most effective for filters having a few state variables, such as IIR filters and the wave- digital filters. For this reason, bit- serial techniques are less frequently applied to FIR structures, especially when the filter length is relatively long "
  • the present invention applies optimization techniques for reducing the area in large sized coefficients by applying a number of optimizations in F--R/LIR filter structures
  • Y(n) c(0) X(n) + c(l) X(n-l) + c(2) X(n-2) + c(n) X(0)
  • Y(n) 5 X(n) + 14 X(n-l) + 25 X(n-2) + 30 X(n-3) + 25 X(n-4) +14 X(n-5) + 5
  • Figure 5" of the drawings shows FIR filter structure of implementation 2. This figure illustrates the realization of FIR filter represented by "Equation 1" .
  • Y(z) X(z)[5*(l+Z- 6 ) + 14*(2r 1 +2T 5 ) + 25*(Z- 2 +2T 4 ) + 30 * Z '3 ] (EQ 2)
  • each column represents a coefficient value.
  • the [T] elements shown as Tl_l to Tljm in columnl, defines connectivity with line SI.
  • [T] elements shown as Tn_l to Tn_m in column n, defines the connectivity with line Sn.
  • T2_m Tn_l to Tn_m is determined by coefficient value.
  • the number of [T] element in a column is dete ⁇ nined.
  • the number of serial adders/subtractor [SA/ SS] in columns is represented as (SA1_1 to SAl_m,SA2_l to SA2_m SAn_l to SAn_m). The presence of one of these elements is again defined by the coefficient value.
  • the [T] elements are arranged in shift register form.
  • the input to first [T] element is connected to one of the S line. While the input to [SA/ SS] is " connected from input S* and or one of the output of [T] elements of shift register, depending on the coefficient value.
  • SAe_l to SAe_n-l elements the addition/subtraction of [SA/SS] of all the coefficient terms depicted in columns is done .
  • the final output is the output of last addition/subtraction[SA SS].
  • the [T] elements are not share-able and also the [SA] in each column are also not share-able. Thus limited minimization is possible in this structure.
  • the device reduces the hardware of the coefficient block [A] by having share-able elements in coefficient block [A], even in the implementation where the coefficient lines CLin_0, CLin_l, are not commonly connected (shown as architecture [A]).
  • the device of the present invention reduces the area by approximately 30- 50% of "Figure 7" of the drawings by reducing the number of components.
  • the optimization techniques are illustrated mathematically and towards the end of this section generalized equation and structure of the device is presented. Accordingly, the present invention (Fig.
  • the device comprising of architecture [A] with hardware sharing techniques and optimization applied to this architecture, the said architecture [A] is connected to coefficient lines CLin_0, CLin_l CLin_n and/or BLin_0, BLin_l,....BLin_n coming from block [E] and/or [F], to be connected to perform filtering operation or a mathematical computing operation with optimization in hardware and provides a zero latency output, the said architecture [A] has serial input bit lines as SI, S2, Sn. [where n represents the number of coefficients of the filter] and the addition terms of the equation [(a0*Sl+b0*S2+....+k0*Sn),
  • Block [B] is a combinational-sequential block consisting of serial adders (SA) & serial subtracters (SS) elements, the connection of elements (SA/SS) to SI,
  • SA/SS elements are arranged in matrix form SA0_0 to SA0_n in bit position 0 and SA1_1 to SAl_n in bit position 1 and similarly SAm_l to
  • SI, S2, S3, S4 represents four inputs.
  • the primary additions are Mone using serial adders SA(1), SA(3), SA(5), SA(8), SA(l l) representing addition of terms S1+S3, S2+S4, S1+S2, S2+S3, S3+S4. While the secondary and tertiary additions are done using the adders SA(6), SA(9), SA(10),
  • the equation defines the bit position as BITO to BIT4, which is the position of "multiplication by power of two", (e.g BITO represents multiplication by 2°).
  • BITO position addition of S3+S4 is performed and the output is terminated at T(l).
  • the output of T(l) defines the next bit position BIT1, which performs addition of S2+S3+S4 and output of T(l) by using the [SA].
  • the output of this addition is again terminated at T(2).
  • the structure is repeated in next BIT positions.
  • the final addition of BIT position BIT4 gives the output of the coefficient block [A].
  • FIG. 10 Implementation of hardware is shown in " Figure 10", wherein the input line SI to S4 represent the lines connected to delay block [Z '1 ] through coefficient line Clin_0 to CLin_6 depicted in " Figure 6" of the drawings.
  • the Lines SI to S4 are connected to block [B] for performing the serial addition/ subtraction, for which [SA], [SS] elements are used within blockfB].
  • the input to [B] block is connected to line SI to S4 and also from [T] elements as would be explained later in this section.
  • the output of each block [B] is terminated with [T] element, which represents the block [B] output being multiplied by "a factor of 2".
  • Each [T] elements defines bit position marked as BIT1, BIT2, BIT3, BIT4.
  • the output b_l of block [B] which is at bit position BITO is fed to the input of the T(l), in turn the output line t_l of element [T(l)] is fed to next section of block[B].
  • the connectivity is done in similar fashion for other [T] blocks.
  • all addition/subtraction in block [B] defines a bit position before getting multiplied by "a factor of 2"and changing to next bit position.
  • the block [B] at final bit position represents the output of the coefficient block [B].
  • Figure 11 shows the implementation of the structure, wherein the input line SI to S4 represent the lines, connected to delay block [Z *1 ] through coefficient lines Clin_0 to CLin_6 depicted in " Figure 6" of the drawings.
  • the Lines SI to S4 are connected to block [B] for performing the serial addition/subtraction, for which (SA), (SS) elements are used within block[B].
  • the input to [B] block is from line SI to S4 and also from [T] elements.
  • the output of each block [B] is terminated with a [T] block, which represents the block [B] output being multiplied by factor of 2.
  • the optimization reduces area for realization of negative coefficient.
  • This optimization is also efficient realization of coefficients having value close to power of two. Further sculpture-r-imization is possible by taking subtraction as common factor and using addition instead of subtraction wherever possible in the realization. This result in an improvement in area, due to the fact that area for a subtractor is more than the area of an adder.
  • the invention provides an area efficient realization of filter coefficient block[A] applicable to filters devices such as FIR, -TR and other filter structures).
  • This architecture is also applicable to combinational and sequential logic consisting of adder, subtracters, multipliers and flip flop [T].
  • This architecture is realized using the elements serial adders (SA), serial subtraction (SS) and flip-flop [T].
  • SA serial adders
  • SS serial subtraction
  • T flip-flop
  • connection of elements (SA/SS) to SI, S2....Sn lines and interconnection of the elements (SA, SS) depend on the value of coefficients [This is because the value of coefficient determines the value of aO, al, etc. and hence it defines the interconnections between them].
  • the output of each block [B] is multiplied by two using [T] elements ⁇ The elements T[l], T[2], T[m] are used for multiplication by factor of 2 ⁇ .
  • the number of T elements depends on the size of maximum coefficient and is share-able for all the coefficient in the coefficient block [A].
  • the final outputs of all the blocks [B] are teiminated at unit delay elements [T] (connected through b_l, b_2,....b_m).
  • flip-flops [T] of all the coefficient are share-able and the number of flip-flops [T] are limited to the coefficient which has the maximum value. Also optimization can be applied in block [D]. The gain in area when compared with the existing design is illustrated below.
  • optimizations in block [D] are referred to as optimization (b), optimization (c) and shown in " Figure 11", “ Figure 12" of the drawings.
  • the invention while in use results in an area improvement of 30-50% of the coefficient block design or combinational logic consisting of adders, subtractor, multiplier and unit delays [T].
  • the present invention is most economical in terms of area of coeff i cient block/ architecture.
  • the present invention provides an area improvement of 30- 50% ) of the coefficient block design or combinational logic consisting of adders, subtractor, multiplier and unit delays.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computing Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computer Hardware Design (AREA)
  • Mathematical Physics (AREA)
  • Complex Calculations (AREA)

Abstract

The invention relates to area efficient realization of coefficient block [A] or architecture [A] with hardware sharing techniques and optimizations applied to this block. The block [A] is connected to coefficient lines CLin_0, CLin_1...CLin_n and BLin_0, BLin_1,... BLin_n coming from block [E] and/or [F], to be connected to perform filtering operation or a mathematical computing operation with optimization in hardware and provides a zero latency output. The invention also gives the area minimal realization of digital filters based on coefficient block [A], when operated in bit serial fashion. The optimization techniques and structure of the present invention are good for linear digital filters typically a finite impulse response (FIR) filter, infinite impulse response filter (IIR) and for other filters and applications based on combinational logic consisting of delay element (T), multiplier (M), adder (SA) and subtractor (SS).

Description

AREA EFFICIENT REALIZATION OF COEFFICIENT ARCHITECTURE FOR BIT-SERIAL FIR, IIR FILTERS AND COMBINATIONAL/SEQUENTIAL LOGIC STRUCTURE WITH ZERO LATENCY CLOCK OUTPUT
FIELD OF THE INVENTION
The invention relates to area efficient realization of coefficient block [A] or architecture [A] with hardware sharing techniques and optimizations applied to this block. The block [A] is connected to coefficient lines CLin_0,
CLin_l CLin_n and BLin_0, BLin_l,....BLin_n co ing from block [E] and/or
[F], to be connected to perform filtering operation or a mathematical computing operation with optimization in hardware and provides a zero latency output. The invention also gives the area minimal realization of digital filters based on coefficient block[A], when operated in bit serial fashion. The optimization techniques and structure of the present invention are good for linear digital filters typically a finite impulse response(F-- ) filter, infinite impulse response filter(IIR) and for other filters and applications based on combinational logic consisting of delay element(T), multiplier(M), adder(SA) and subtractor (SS).
Brief description of the accompanying drawings
In the accompanying drawings:
Figure 1 shows the field of invention, applications of the device
Figure 2 shows the symbol of components used in device.
Figure 3 shows the description of the components used in device.
Figure 4 shows the bit serial FIR filter implementations.
Figure 5 shows an example of FIR filter.
Figure 6 shows one of the known minimization technique due to symmetry .of coefficient. Figure 7 shows the structure of existing/known implementation technique for example FIR filter.
Figure 8 shows the generalized structure of existing/known implementation technique of coefficient block.
Figure 9 shows a realization of the coefficient block for coefficient value close to power of two.
Figure 10 shows the optimization (a) for realizing the coefficient block for example FIR filter, of the present invention.
Figure 11 shows the optimization (b) for realizing the coefficient block for example FIR filter, of the present invention.
Figure 12 shows the optimization (c) for realizing the coefficient block for example FIR filter, of the present invention.
Figure 13 shows the generalized optimized structure, of the present invention.
Details of Elements/symbol used in the description
The basic components symbol used in design are shown in "Figure 2" of the drawings. In addition, explanation and usages of the device are done in the text below & depicted in "Figure 3" and "Figure 4" of the drawings. Unit delay (T)
It is one bit delay element. It also performs function of a multiplier by a factor of 2. [e.g. For the serial input frame (0101011 in binary or 43 in integer representation), the output of this block is (01010110 in binary or 86 in integer representation). This element is usually a Flip-flop (D Flip-flop, J-K Flip-flop etc.). Full adder (FA)
It performs binary addition. The inputs to this element are A, B, Cin (Carryin) while the outputs are Z and Cout (Carryout). The truth table for full adder functionality is shown in "Figure 3" of the drawings. Full subtractor (FS)
It performs binary subtraction. The inputs to this element are A, B, Cin (Carryin) while the outputs are Z and Cout (Carry out). The truth table for full subtractor functionality is shown in "Figure 3"of the drawings. Serial adder (SA) and Serial Subtractor (SS)
It performs addition/subtraction of two serial frame, xl(nT), x2(nT) to generate output y(nT) represented as xl(nT)+x2(nT) or xl(nT)-x2(nT) . The serial adder (or subtractor) is implemented using a full adder (or subtractor) with a Flip-Flop as shown in "Figure 3" of the drawings. The output Cout of [FA/FS] is delayed using the [T] element and is applied to Cin line of [FA/ FS]. This enables the [FA FS] and [T] together to function as serial adder (SA/SS), where A, B are the inputs to this element and Z is the output, (e.g of serial addition is as follows, if xl(nT) = 0110 (6 in integer) and x2(nT) = 0111 (7 in integer). Then y(nT) = 01101 (13 in integer representation). Serial Multiplier (M)
It multiplies two serial input frame X(nT) and m. The output is a function represented as Y(nT) = X(nT) * m. A serial coefficient multiplier(M) can be implemented by shift register using [T] elements and adder element [SA] (One shift means multiply by factor of 2). As shown in "Figure 3" of the drawings, the multiplier is formed by adding the outputs corresponding to ones in the binary representation of the coefficient. Delay (Z"1)
" Delay by one frame of data is done by shift register (series of Flip-flops (T) connected to store and shift the input frame). The number of Unit delay (T) in one delay element is equal to the frame size of the input. PRIOR ART OR EXISTING IMPLEMENTATION OF FILTER
The following description discusses the elements used for implementation of design and the existing implementations for digital filters. The proposed minimization is extendable to other applications such as Digital Signal Processing field and Digital designs.
From here onwards, all the illustration would be done with the preferred embodiment namely FIR filter which is extendable to other filters as described earlier. "Figure 4" shows the existing structure of bit serial FLR filter with coefficient lines CLin_0, CLin_l, CLin_n and the coefficient block [A] having the coefficients c(0), c(l), c(2),....c(n). The coefficient block is connected to delay element [Z'1] and serial adders [SA] to form filter structure. Stating the FIR filter equation in time and frequency domain
Y(n) = c(0) X(n) + c(l) X(n-1) + c(2) X(n-2) + c(n) X(0)
Y(z) = X(z) [c(0) + c(l) Z"1 + c(2) Z"2 + c(3) Z"3 + c(4) ZA + c(5) Z"5 + c(6)
Z'6+ +c(n) Z"n] where X, Y are the input and output respectively and c(0), c(l) c(n) represent the coefficients value which defines the characteristics of the filter and each delay [Z"1] block represents one sample delay. The filter equation can be implemented in two ways as shown in "Figure 4" of the drawings
In implementation 1, coefficient lines CLin_0, CLin_l, CLin_n are common and connected to input X[n]. The output lines CLout_0, CLout_l, CLout_n are connected to block [E], consisting of delay element [Z"1] and serial adders [SA] elements. The structure makes easy realization of share-able multiplier in the coefficient block [A]. An example of share-able multiplier with coefficient values 3,11 is illustrated in "Figure 4". The realization of these coefficient separately would require 4[T], 3[SA] elements. By virtue of CLin_0, CLin_l,... being common, the hardware is realized using 3[T], 2[SA] elements. Another feature of the structure is that the structure inherently requires more storage area, represented by [Z"1], as compared to implementation 2, since the storage is done after the multiplication. For input frame of n bit and coefficient of size m bit, the storage area of each delay element [Z"1] is (m+n). The total storage space of the delay elements is (m+n) * (number of coefficients -1).
In implementation 2, the coefficient line CLin_0, CLin_l, are not common. By virtue of connectivity of different input lines to the coefficient elements [c(0), c(l) ], the realization of coefficient block [A] using share-able elements is not present. Another feature of this structure is that it inherently requires lesser storage space, represented as [Z"1], as unlike in previous implementation, here the storage is done before multiplication. For input frame of m bit and coefficient of size n bit, the storage area of each delay element [Z'1] is (m). The total storage space is (m) * (number of coefficients -1).
The invention reduces the hardware of the coefficient block [A] by having shareable elements in coefficients, even if the coefficient lines CLin_0, CLin_l, are not commonly connected. For existing configuration as shown in "Figure 7" and "Figure 8", the share-ability of hardware in block [A] is a limitation.
Also, as described in previous section, implementation 2 is area efficient with "respect to implementationl due to reduced delay elements size. Over and above this by having share-able multiplier or reduced coefficient block [A], which are the key features of the invention, implementation 2 becomes still more area-efficient. This reduction is extendable to other filter based on coefficient block [A] as stated in the first section. The present invention operates on integer valued coefficient. Further, to quote Norsworthy and Crochiere (Delta-Sigma Data Converters LEEE press pp-435, copyright 1997)
"Bit-serial architecture reduce the interprocessor communication down to 1 bit. Generally the number of processors is very large, but because each processor is so small, the overall economy is very high. Bit serial architectures are usually most effective for filters having a few state variables, such as IIR filters and the wave- digital filters. For this reason, bit- serial techniques are less frequently applied to FIR structures, especially when the filter length is relatively long "
However, the present invention applies optimization techniques for reducing the area in large sized coefficients by applying a number of optimizations in F--R/LIR filter structures
To elaborate the applicant's optimization techniques, consider an FIR filter with symmetrical coefficient as 5, 14, 25, 30, 25, 14, and 5. Though, the size of the coefficients in this example is small, it is enough to elaborate the minimization proposals. Stating the FIR filter equation in time and frequency domain
Y(n) = c(0) X(n) + c(l) X(n-l) + c(2) X(n-2) + c(n) X(0)
Y(z) = X(z) [c(0) + c(l) Z"1 + c(2) Z2 + c(3) Z"3 + c(4) Z-4 + c(5) Z5 + c(6)
Z"6+ +c(n) Z'n] where X, Y are the input and output respectively and c(0), c(l) c(n) represent
-the coefficients value.
Using the coefficient values in above equation
Y(n) = 5 X(n) + 14 X(n-l) + 25 X(n-2) + 30 X(n-3) + 25 X(n-4) +14 X(n-5) + 5
X(n-6)
Y(z) = X(z) [5 + 14 Z + 25Z-2 + 30 Z-3 + 25 Z4 + 14 Z5 + 5 Z'6] (EQ 1) The Existing Method and Minimization
"Figure 5" of the drawings shows FIR filter structure of implementation 2. This figure illustrates the realization of FIR filter represented by "Equation 1" .
In one of the known optimization technique, is taken advantage of the symmetry in the coefficients. The streams which have to be multiplied with the same coefficients can be added first and then multiplied. For a large filter structure, this leads to a reduction by 45% in the coefficient block, (see "Figure 6" of the accompanying drawings)
This is done by rest cturing the equation as under:
Y(z) = X(z)[5*(l+Z-6) + 14*(2r1+2T5) + 25*(Z-2+2T4) + 30 * Z'3] (EQ 2)
For the rest of the optimization proposals it will be talking about only the multiplier adder series which is shown in the dotted box referred to as coefficient block [A]. "Figure 7" of the drawings shows the traditional way of implementation of the example structure for block [A], wherein SI to S4 represent the lines connected to delay block [Z"1] through line CLin_0 to CLin_6 depicted in "Figure 6" of the drawings. The Lines SI to S4 are separately connected to [T] element for performing a multiplication by "a factor of 2" and (SA) is being used to perform serial addition of data. This represents the multiplier less realization of filter coefficient block [A)] where the property of flip-flop (T) as multiplier of two is used. Mathematically, the restructured equation according to the structure is stated as
Y(nT)=(4+l)Sl + (8+4+2)S2 + (16+8+l)S3 + (16+8+4+2)S4 (EQ 3)
In this implementation, SI, S2, S3, S4 lines are not commonly connected. Hence this restricts to achieve a share-able hardware in coefficient block [A]. Thus all the function/operations of this block represent unique hardware. The elements required by the terms are listed as
First term = 2[T], 1[SA] elements
Second term = 3[T], 2[SA] elements
Third term = 4[T], 2[SA] elements
Fourth term = 4[T], 3[SA] elements
Final addition of all the four term would require 3[SA] elements.
The generalized structure of "The Existing Method and Minimization" is depicted in "Figure 8". In the structure, each column represents a coefficient value. The [T] elements, shown as Tl_l to Tljm in columnl, defines connectivity with line SI. In similar fashion, [T] elements, shown as Tn_l to Tn_m in column n, defines the connectivity with line Sn.
The presence of one of the elements in columns 1 to n (i.e Tl_l to Tl_m, T2_l to
T2_m Tn_l to Tn_m) is determined by coefficient value. Thus depending on coefficient value on lines SI to Sn, the number of [T] element in a column is deteπnined. Also the number of serial adders/subtractor [SA/ SS] in columns is represented as (SA1_1 to SAl_m,SA2_l to SA2_m SAn_l to SAn_m). The presence of one of these elements is again defined by the coefficient value.
In the structure, the [T] elements are arranged in shift register form. The input to first [T] element is connected to one of the S line. While the input to [SA/ SS] is "connected from input S* and or one of the output of [T] elements of shift register, depending on the coefficient value. Finally, using SAe_l to SAe_n-l elements, the addition/subtraction of [SA/SS] of all the coefficient terms depicted in columns is done . The final output is the output of last addition/subtraction[SA SS]. Among the lines SI to Sn, the [T] elements are not share-able and also the [SA] in each column are also not share-able. Thus limited minimization is possible in this structure.
DETAILED DESCRIPTION OF THE INVENTION
The device reduces the hardware of the coefficient block [A] by having share-able elements in coefficient block [A], even in the implementation where the coefficient lines CLin_0, CLin_l, are not commonly connected (shown as architecture [A]).
This reduced hardware in coefficient block when applied in implementation-. ("Figure 4") makes it still more area- efficient. This reduction is extendable to other filter based on coefficient block [A], as stated in the first section.
The device of the present invention reduces the area by approximately 30- 50% of "Figure 7" of the drawings by reducing the number of components. The optimization techniques are illustrated mathematically and towards the end of this section generalized equation and structure of the device is presented. Accordingly, the present invention (Fig. 13) provides a device for providing an area efficient realization of coefficient, said device comprising of architecture [A] with hardware sharing techniques and optimization applied to this architecture, the said architecture [A] is connected to coefficient lines CLin_0, CLin_l CLin_n and/or BLin_0, BLin_l,....BLin_n coming from block [E] and/or [F], to be connected to perform filtering operation or a mathematical computing operation with optimization in hardware and provides a zero latency output, the said architecture [A] has serial input bit lines as SI, S2, Sn. [where n represents the number of coefficients of the filter] and the addition terms of the equation [(a0*Sl+b0*S2+....+k0*Sn),
(al*Sl+bl*S2+ +kl*Sn) (am*Sl+bm*S2+ +km*Sn)] are represented as blocks [B] , the values of aO, bO, etc. are represented as [(+ /-)1 or 0], the said Block [B] is a combinational-sequential block consisting of serial adders (SA) & serial subtracters (SS) elements, the connection of elements (SA/SS) to SI,
S2....Sn lines and interconnection of the elements (SA, SS) depend on the value of coefficients, the SA/SS elements are arranged in matrix form SA0_0 to SA0_n in bit position 0 and SA1_1 to SAl_n in bit position 1 and similarly SAm_l to
SAm_n in bit position m, the presence of one of these elements is defined by coefficient value, the output of each block [B] is connected to [T] elements through line b_l, b_2,....b_m, the number of T elements depends on the size of maximum coefficient and is share-able for all the coefficient in the coefficient architecture [A], the output of element [T] is connected to one of the inputs of combinational logic of block [B] of next bit position (i.e connected to input of element (SA or SS) of block [B] depending upon the sign value+/-), Lines t_l, t_2, t_m are used to mark the interconnections from cluster [C] to [B], in the said structure [A], all the elements in block [B] are clustered together as block [D] and all the unit delay elements {T[l], T[2] T[m]} are clustered together in block [C], thereby separating the combinational-sequential and sequential logic, while the sequential elements [T] of block [C] are common for all the coefficients and are share-able and positioned at end position of each block [B], the Block [D] has combinational-sequential element block [B] which are essentially SA, SS, In the structure the hardware within block [B] are share- able across various [B] blocks and also within block [D]. The final output is taken from the output of the elements of last bit position.
An embodiment of the invention or Optimization (a)
Continuing the same example of FIR filter & using "Equation 3" of previous section. y(nT)= 5 * SI + 14 * S2 + 25 * S3 + 30 * S4
Y(nT)=(4+l)Sl + (8+4+2)S2 + (16+8+l)S3 + (16+8+4+2)S4 we proceed to share the shift registers (multiply by 2) of the design.
=(S3+S4)* 16+(S2+S3+S4)*8+(S1+S2+S4)*4+(S2+S4)*2+(S1+S3)
=(S1+S3)+2*(S2+S4+2*(S1+S2+S4+2*(S2+S3+S4+2*(S3+S4)))) (EQ 4)
The implementation flow for this equation is presented below this text paragraph and the hardware implementation is shown in "Figure 10" of the drawings. In the flow of implementation, SI, S2, S3, S4 represents four inputs. The primary additions are Mone using serial adders SA(1), SA(3), SA(5), SA(8), SA(l l) representing addition of terms S1+S3, S2+S4, S1+S2, S2+S3, S3+S4. While the secondary and tertiary additions are done using the adders SA(6), SA(9), SA(10),
SA(7), SA(4), SA(2). The multiplication by factor of two is done using the elements T(l), T(2), T(3), T(4).
Implementation flow of equation {optimization(a)}
Figure imgf000013_0001
As shown in the above implementation flowchart, the equation defines the bit position as BITO to BIT4, which is the position of "multiplication by power of two", (e.g BITO represents multiplication by 2°). At BITO position addition of S3+S4 is performed and the output is terminated at T(l). The output of T(l) defines the next bit position BIT1, which performs addition of S2+S3+S4 and output of T(l) by using the [SA]. The output of this addition is again terminated at T(2). The structure is repeated in next BIT positions. The final addition of BIT position BIT4 gives the output of the coefficient block [A].
Implementation of hardware is shown in "Figure 10", wherein the input line SI to S4 represent the lines connected to delay block [Z'1] through coefficient line Clin_0 to CLin_6 depicted in "Figure 6" of the drawings. The Lines SI to S4 are connected to block [B] for performing the serial addition/ subtraction, for which [SA], [SS] elements are used within blockfB]. The input to [B] block is connected to line SI to S4 and also from [T] elements as would be explained later in this section. The output of each block [B] is terminated with [T] element, which represents the block [B] output being multiplied by "a factor of 2". Each [T] elements defines bit position marked as BIT1, BIT2, BIT3, BIT4. The output b_l of block [B] which is at bit position BITO is fed to the input of the T(l), in turn the output line t_l of element [T(l)] is fed to next section of block[B]. The connectivity is done in similar fashion for other [T] blocks. Thus all addition/subtraction in block [B] defines a bit position before getting multiplied by "a factor of 2"and changing to next bit position. The block [B] at final bit position represents the output of the coefficient block [B].
In the structure, all [T] elements are represented as block[C] wherein the flip- flop [T] representing multiplication by "a factor of 2", is share-able among various coefficient values and their number is determined by maximum coefficient value. This is in contrast to "Figure 7" of existing structure where the elements are not share-able between SI to S4 lines.
The number of flip-flops (T) in "Figure 7" is 13 vs. number of flip-flops(T) in current proposal is 4. Also for both the implementation, the number of the one-bit serial adders (SA) remains the same [11 in each case]. In present minimization, approximate area calculations is = 11 serial adder + 4 T = 26 Units, whereas the area after previous minimization is 11 serial adder + 13 T = 35 units, (assuming 1 Unit = 1 FA = 2HA = IT and serial adder/serial subtractor (SA/SS) = 2 Units). This is approximately 26%saving in area in "Figure 10" as compared to "Figure 7".
For filter having large size coefficient, this leads to a drastic reduction in the area (30-50% of the coefficient block).
Another embodiment of the invention or Optimization (b)
This optimization reduces the hardware of block [D] which essentially consists of
(SA) and (SS) elements. Beginning with "Equation 4" and finding out the common additive factors
Al = S2+S4
A2 = S3+S4
The "Equation 4" can be further reduced as y(nT) = (Sl+S3)+2*(Al+2*(Sl+Al+2*(S2+A2+2*A2))) (EQ 5)
The flow of implementation of equation is illustrated below and is self explanatory. Here SI, S2, S3, S4 represents four inputs. The primary addition is done using serial adders SA(1), SA(3), SA(9) representing addition of terms
S1+S3, S2+S4, S3+S4. While the secondary and tertiary addition is done using the adders SA(5), SA(7), SA(8), SA(6), SA(4), SA(2). The multiplication by factor of two is done using the elements T(l), T(2), T(3), T(4).
Implementation flow of equation {optimization (b)}
B1T4 BIT3 BIT2 BITl BITO
=
Figure imgf000016_0001
"Figure 11 " shows the implementation of the structure, wherein the input line SI to S4 represent the lines, connected to delay block [Z*1] through coefficient lines Clin_0 to CLin_6 depicted in "Figure 6" of the drawings. The Lines SI to S4 are connected to block [B] for performing the serial addition/subtraction, for which (SA), (SS) elements are used within block[B]. The input to [B] block is from line SI to S4 and also from [T] elements. The output of each block [B] is terminated with a [T] block, which represents the block [B] output being multiplied by factor of 2. The output b_l of block [B] which is at bit position BITO, is fed to the input of the T(l), in turn the output t_l of [T(l)] is fed to next section of block[B]. Thus all addition defines a bit position before getting multiplied by factor of 2. All such [T] termination is represented by block[C].
The optimizations in reducing the hardware of block [D] are done. The output b_l representing the bit position BITO and addition term A2, is connected to T[l] and also fed to the next block [B], hence reducing the adder count by 1. Also the output of adders SA(3) of block [B] in bit position BIT3, is fed at two points. One to the input of adders S A(4) which eventually terminates at [T4] element and other to the input of adder SA(5), hence reducing the adder count further by 1. Note how Al and A2 are shared in the Structure. Comparing the hardware implementation of "Figure 10" and "Figure 11", the number of adders are minimized by having common adders Al, A2. This optimization is purely dependent on finding common addition terms among coefficients.
In present minimization, approximate area calculations is = 9 serial adder (SA) + 4 (T) = 22 Units, whereas the area of the existing minimization of "Figure 7" is 11 (SA)+ 13 (T) = 35 units, (assuming 1 Unit = 1 FA = 2HA = IT and serial adder (SA) = 2 Units). Thus compared to the existing minimization, the "Optimization (a)" and "Optimization (b)" , combined has resulted in 37% saving in area (13 / 35 * 100). "Optimization (b)" is an improvement of 15% in area (of the coefficient block) over "Optimization (a)"
Yet another embodiment of the invention or Optimization (c)
In realization of block [D], further optimization is done by realizing the coefficient value using subtraction instead of addition. This is good for numbers which have values closer to power of 2. (e.g. for realization of coefficient value 63, the realization (63=64 -1) is better than (63 = 32+16+8+4+2+1). In first case the number of subtractor is 1 while in second case the number of adders are 5. To illustrate this by an example, consider the coefficient values as 5, 25, -48, -63). Writing the FIR equation using these coefficient values.
Arranging the terms with 63 as (32+16+8+4+2+1) y(nT)= 5 * SI + 25 * S2 - 48 * S3 - 63 * S4
= (1+4)*S1 + (16+8+l)*S2 - (32+16)*S3 - (32+16+8+4+2+l)*S4
=(Sl+S2-S4)+2*(-S4+2*(Sl-S4+2*(S2-S4+2*(S2-S4- (S3+2(S3+S4)))))) (EQ 6)
Alternately arranging the terms with 63 as (64 -1), the first equation reduces
=(1+4)*S1 + (16+8+l)*S2 - (32+16)*S3 + (1-64)*S4
=(S1+S2+S4)+2*(2*(S1+2*(S2+2*(S2-S3-2*(S3+2*S4))))) (EQ 7)
The realization of "Equation 6" and "Equation 7" is shown in "Figure 9" and "Figure 12" respectively.
In these realizations, the number of [T] elements is one more in "Equation 7" due to presence of term "64". However, the number of adders are less in the structure represented by "Equation 7" than by "Equation 6" . This is because the number of adders are less in former case.Comparing the area of the two realization, From "Equation 6" , Area obtained is 5 T+6 SA+6 SS = 29 Units. While from "Equation 7" , representing optimization(c), results an area calculation of = 6T + 6SA + 2SS = 22 Units, (assuming 1 Unit = 1 FA = 2HA = IT and SA=SS= 2 Units). This is an improvement by 24% in reducing area of coefficient block for the current example.
Thus, the optimization reduces area for realization of negative coefficient. This optimization is also efficient realization of coefficients having value close to power of two. Further irii-r-imization is possible by taking subtraction as common factor and using addition instead of subtraction wherever possible in the realization. This result in an improvement in area, due to the fact that area for a subtractor is more than the area of an adder.
GENERALIZED STRUCTURE OF THE -INVENTION
The invention provides an area efficient realization of filter coefficient block[A] applicable to filters devices such as FIR, -TR and other filter structures). This architecture is also applicable to combinational and sequential logic consisting of adder, subtracters, multipliers and flip flop [T]. This architecture is realized using the elements serial adders (SA), serial subtraction (SS) and flip-flop [T]. Generalized structure of the present invention is depicted in "Figure 13". The generalized equation of the structure is also calculated here.
Beginning with the generalized equation of FIR filter coefficient block(A) y(nT) = a * SI + b * S2 + c* S3 + k * Sn (1) where a, b,....k represents filter coefficients. SI, S2 represents bit lines corresponding to the coefficients.
Now, representing each coefficient as addition of terms arranged in power of two and applying it to the equation: y(nT) = (2m*am + 21*al+2°*a0) * SI + (2m*bm + 2'*bl+2°*b0) * S2 +
(2m*cm + 21*cl+2°*c0) * S3+ +(2m *km + 2"1*kl+2°*k0) * Sn
Further, taking "2" as common factor we get the generalized equation for architecture as:
Y(nT) = (aO*Sl +bO*S2+....+kO*Sn)
+ 2λ ( (al *Sl+bl*S2+....+kl*Sn) +
21 ((a2*Sl+b2*S2+...+k2*Sn)+
21((a3*Sl+b3*S2+...+k3*Sn)+ +
21 ((am*Sl+bm*S2+ +km*Sn))))) where aO, al, am and bO, bl,...bm and kO, kl, km represents the sign of coefficients [i.e they have value (+ / - 1) or 0]. The architecture realization in "Figure 13" is done using the sequential elements like unit delays [T] and combinational elements such as serial adder (SA) and serial subtractor (SS). The highlights of the architecture are:
1) A common share-able [T] elements for all the coefficients. The maximum number of [T] element is equal to next integer value of "log of the maximum value of coefficient" in the coefficient block [A]
2) Area-optimizations in reducing the combinational logic [D] (i.e optimizations applied on serial adders(SA), serial subtractor (SS) as stated in previous section).
In "Figure 13", the input serial data is present on bit line SI, S2, Sn. [where n represents the number of coefficients of the filter] The addition terms of the equation.[(aO*Sl+bO*S2+....+kO*Sn),(al*Sl+bl *S2+ +kl*Sn) (am*Sl+b m*S2+ +km*Sn)] are represented as blocks [B]. Block [B] is a combinational- sequential block consisting of serial adders (SA) and serial subtracters (SS) elements. Since the values aO, bO, etc. represents value [(+ / -)1 or 0]. The connection of elements (SA/SS) to SI, S2....Sn lines and interconnection of the elements (SA, SS) depend on the value of coefficients [This is because the value of coefficient determines the value of aO, al, etc. and hence it defines the interconnections between them]. The output of each block [B] is multiplied by two using [T] elements {The elements T[l], T[2], T[m] are used for multiplication by factor of 2}. The number of T elements depends on the size of maximum coefficient and is share-able for all the coefficient in the coefficient block [A]. Thus in the structure the final outputs of all the blocks [B] are teiminated at unit delay elements [T] (connected through b_l, b_2,....b_m).
In the structure, all the elements [B] are clustered together as [D] and all the unit delay elements {T[l], T[2] T[m]} are clustered together in [C]. The sequential
[C] and combinational-sequential logic [D] are quite separated in this architecture. The input of the unit delay element [T] is final output of block [B] and the output of element [T] is connected to the one of the inputs of combinational logic of block [B] of next bit position (i.e connected to input of element (SA or SS) of block [B] depending upon the sign value+/-). The interconnections from cluster [C] to [B] is represented as t_l, t_2, t_m. The bit positions of serial data frame are marked as
BIT0, BIT1, BITm.
In the generalized structure, flip-flops [T] of all the coefficient are share-able and the number of flip-flops [T] are limited to the coefficient which has the maximum value. Also optimization can be applied in block [D]. The gain in area when compared with the existing design is illustrated below.
Hardware reduction in block [C]
Before beginning to prove the statement, we proceed to formularize the calculation of number of flip-flop (T) for structure of "The Existing Method & Minimization" in "Figure 7" and "Figure 8" of the drawings. The number of flip-flops in the coefficient block depends on the size of all the coefficients. The approximate and pessimistic formula for calculation of total flip-flops (T) in coeff. block is [= average size of coefficient * number of coefficient], where average size of coefficient can be calculated pessimistically as (Maximum coefficient size / 2). (Refer "The Existing Method and Minimization", "Figure 8" of the drawings). Applying this formula, to example of "Figure 7" for verification, where coefficient (5, 14, 25, 30) are represented in 4,5,6,6 bits (using signed notation). According to formula, average size of coefficient is (6 / 2) = 3 and total number of flip-flop = 3 * 4 = 12. This is pessimistic as compared to the implementation (where total number of flip-flops are 13, refer "Figure 7").
Similarly, the approximate formula for calculation of total adders (SA) in coeff. block for "The Existing Method and -v--iιτJ-mization" , and "Detailed Description of the Invention" in "Figure 8" and "Figure 13" is [=adders per coefficient * number of coefficient]. Adders per coefficient solely depend on the value of coefficient. Assuming number of adders as (=number of coefficients * maximum coefficient size / 2).
Now, as an example, the Applicants provide herebelow use of the above mentioned formulae of previous two paragraphs in filter of 20 coefficient. Assume the maximum coefficient value is represented in 16 bits (e.g. maximum coefficient value is +32767 or -32768 in 2's complement representation). Average size of the coefficient approximated by the formula is (16/2=8 bit). In "The Existing Method and Minimization", the total number of flip-flop (T) required for implementation is 8 * 20 = 160. In contrast to this "Detailed Description of the Invention" would require only 16 Flip-Flops (The number of flip-flops of all the coefficient are share- able and are limited to the coefficient which has the maximum value). Assuming in worst case that there is no optimization of adders, the number of adders in both the cases are same and are equal to 8 * 20 = 160. (Refer "Figure 8" and "Figure 13" of the drawings)
Area calculation for "The Existing Method and Minimization" , "Figure 8" of the drawings is 160 T +160 SA = 480. Area calculation for "Detailed Description of the Invention" , "Figure 13" is 16 T +160 SA = 336. This is an improvement of 30% [(480-336)/480] over "The Existing Method and Minimization" . (assuming 1 Unit = 1 FA = 2HA = IT and serial adder/serial subtractor (SA/SS) = 2 Units). The area gain by structure could be as high as 50% (of the coefficient block) for big filter where ιτι--nimization of adders and other minimization optimization^), optimization(b) and optimization(c), as discussed earlier, are applied.
The preferred embodiment of the invention is also supported by a real example of filter coefficient device. This is referred to as optimization(a) and shown in "Figure 10" of the drawings and discussed in previous section. Area calculation for "Figure 10" is 11SA+4T=26 units while area of "Figure 7" of the drawings (existing implementation) is 11SA+13T=35 units. This results a gain of 26% in area (of coefficient design block) for this example design as compared to existing implementation, supporting the generalized statement.
Hardware reduction in block [D]
The optimizations in block [D] are referred to as optimization (b), optimization (c) and shown in "Figure 11", "Figure 12" of the drawings. In optimization (b), beside sharing the flip-flop (T) for all the coefficients and sharing the common adders (SA) techniques is done in cluster [D]. This is due to presence of (SA), (SS) in block [D] and separate clustering of elements in block [D] and [C]. This is illustrated using the example of previous section Using Al = S2+S4, A2 = S3+S4 in this example. Total area for this example after the optimization is 9SA + 4T = 22 Units*. This, when combined with optimization(a) results in area-gain of 31% in area (of coefficient design block) for this example design as compared to existing implementation (where area was 35 units).
The optimization(c) as described before can be further applied to cluster [D]. That is beside optimization(a) and optimization(b), the technique of realizing the coefficient value using subtraction (SS) instead of addition(SA) is used here. This saves quite in area when the coefficient value is close to power of 2. (e.g. for realization of coefficient value 63, the realization (63=64 -1) is better than (63 = 32+16+8+4+2+1). In first case the number of subtractor is 1 while in second case the number of adders are 5.). The two cases are illustrated in previous section with example and are shown in "Figure 12" and "Figure 9" of the drawings. The area calculation without optimization and using adders ("Figure 9" of the drawings) = 29 Units*. While optimization applied ("Figure 12" of the drawings) results in area = 22 Units*. This is an improvement by 24% in reducing area of coefficient block for this example.
With all the optimization applied, the invention while in use results in an area improvement of 30-50% of the coefficient block design or combinational logic consisting of adders, subtractor, multiplier and unit delays [T].
*Note that the input to adders in [B] are interchangeable e.g adders SA(5), SA(6) inputs could be interchanged. Also the signals t_l , t_2 etc. can be connected to any input of adders of block [B] of next bit position. **For approximate area calculation following assumption is made (1 Unit of Area = 1 FA = 2HA = IT & SA=SS= 2 Units of Area).
The present invention is most economical in terms of area of coefficient block/ architecture. In fact, the present invention provides an area improvement of 30- 50%) of the coefficient block design or combinational logic consisting of adders, subtractor, multiplier and unit delays.

Claims

CLAEVIS
1. A device for providing an area efficient realization of coefficient, said device comprising of architecture [A] with hardware sharing techniques and optimization applied to this architecture, the said architecture [A] is connected to coefficient lines CLin_0, CLin_l CLin_n and/or BLin_0, BLin_l,....BLin_n coming from block [E] and/or [F], to be connected to perform filtering operation or a mathematical computing operation with optimization in hardware and provides a zero latency output, the said architecture [A] has serial input bit lines as SI,
S2, Sn. [where n represents the number of coefficients of the filter] and the addition terms of the equation [(aO*Sl+bO*S2+....+kO*Sn),
(al *Sl+bl*S2+ +kl *Sn) (am*Sl+bm*S2+ +km*Sn)] are represented as blocks [B] , the values of aO, bO, etc. are represented as [(+ /-)1 or 0], the said
Block [B] is a combinational-sequential block consisting of serial adders (SA) and serial subtracters (SS) elements, the connection of elements (SA/SS) to SI, S2....Sn lines and interconnection of the elements (SA, SS) depend on the value of coefficients, the SA/SS elements are arranged in matrix form SA0_0 to SA0_n in bit position 0 and SA1_1 to SAl_n in bit position 1 and similarly SAm_l to SAm_n in bit position m, the presence of one of these elements is defined by coefficient value, the output of each block [B] is connected to [T] elements through line b_l, b_2,....b_m, the number of T elements depends on the size of maximum coefficient and is share-able for all the coefficient in the coefficient architecture [A], the output of element [T] is connected to one of the inputs of combinational logic of block [B] of next bit position (i.e connected to input of element (SA or SS) of block [B] depending upon the sign value+/-), Lines t_l, t_2, t_m are used to mark the interconnections from cluster [C] to [B], in the said structure [A], all the elements in block [B] are clustered together as block [D] and all the unit delay elements (T[l ], T[2] T[m]} are clustered together in block [C], thereby separating the combinational-sequential and sequential logic, while the sequential elements [T] of block [C] are common for all the coefficients and are share-able and positioned at end position of each block [B], the Block [D] has combinational-sequential element block [B] which are essentially SA, SS, In the structure the hardware wilbin block [B] are share- able across various [B] blocks and also within block [D]. The final output is taken from the output of the elements of last bit position.
2. A device as claimed in claim 1 wherein the area ir-inimal realization of digital filters based on coefficient architecture[A], when operated in bit serial fashion, the structure [A] provides hardware minimization for finite impulse response(FIR) filter, infinite impulse response filter(IIR) and for other applications related to combinational logic consisting of delay element(T), multiplier(M), adder(SA) and subtractor (SS).
3. A device as claimed in claim 1 wherein further optimization technique in cluster [D] is done by using common adders (SA) and common subtractor (SS) in Block [B] and using this shared outputs.
4. A device as claimed in claiml wherein further optimization technique in cluster [D] is done by using subtractor (SS) instead of adders, when the coefficient value is closer to power of two.
5. A device as claimed in claim 1 wherein further optimization technique in cluster [D] is done by minimizing the use of subtractor by taking common subtraction operator and using adder instead.
6. A device as claimed in anyone of the preceding claims wherein when used in implementation 2 of FIR/IIR filter and similar structure of filters, results in quite area efficient realization of filter, this is due to the fact that storage area of delay elements of filter in implementation2 is smaller in size as compared to implementation 1 which is present due to inherent property of the structure of implementation 2, An additional saving in area in filter coefficient realization design is achieved by using the claimed structure of coefficient architecture [A] of "Figure 13"
PCT/SG1998/000081 1998-10-13 1998-10-13 Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output Ceased WO2000022728A1 (en)

Priority Applications (4)

Application Number Priority Date Filing Date Title
PCT/SG1998/000081 WO2000022728A1 (en) 1998-10-13 1998-10-13 Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output
DE69821144T DE69821144T2 (en) 1998-10-13 1998-10-13 AREA EFFICIENT MANUFACTURE OF COEFFICIENT ARCHITECTURE FOR BIT SERIAL FIR, IIR FILTERS AND COMBINATORIAL / SEQUENTIAL LOGICAL STRUCTURE WITHOUT LATENCY
EP98950601A EP1119909B1 (en) 1998-10-13 1998-10-13 Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output
US10/968,822 US20050193046A1 (en) 1998-10-13 2004-10-21 Area efficient realization of coefficient architecture for bit-serial fir, IIR filters and combinational/sequential logic structure with zero latency clock output

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/SG1998/000081 WO2000022728A1 (en) 1998-10-13 1998-10-13 Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US10/968,822 Continuation US20050193046A1 (en) 1998-10-13 2004-10-21 Area efficient realization of coefficient architecture for bit-serial fir, IIR filters and combinational/sequential logic structure with zero latency clock output

Publications (1)

Publication Number Publication Date
WO2000022728A1 true WO2000022728A1 (en) 2000-04-20

Family

ID=20429881

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/SG1998/000081 Ceased WO2000022728A1 (en) 1998-10-13 1998-10-13 Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output

Country Status (4)

Country Link
US (1) US20050193046A1 (en)
EP (1) EP1119909B1 (en)
DE (1) DE69821144T2 (en)
WO (1) WO2000022728A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1650869A1 (en) * 2004-10-20 2006-04-26 STMicroelectronics Pvt. Ltd A device for implementing a sum of products expression

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP4265642B2 (en) * 2006-10-16 2009-05-20 ソニー株式会社 Information processing apparatus and method, recording medium, and program

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO1994023493A1 (en) * 1993-04-05 1994-10-13 Saramaeki Tapio Method and arrangement in a transposed digital fir filter for multiplying a binary input signal with tap coefficients and a method for disigning a transposed digital filter

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
DE69821145T2 (en) * 1998-10-13 2004-09-02 Stmicroelectronics Pte Ltd. AREA EFFICIENT MANUFACTURE OF COEFFICIENT ARCHITECTURE FOR BIT SERIAL FIR, IIR FILTERS AND COMBINATORIAL / SEQUENTIAL LOGICAL STRUCTURE WITHOUT LATENCY

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO1994023493A1 (en) * 1993-04-05 1994-10-13 Saramaeki Tapio Method and arrangement in a transposed digital fir filter for multiplying a binary input signal with tap coefficients and a method for disigning a transposed digital filter

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
DAWOOD ALAM ET AL: "VLSI IMPLEMENTATION OF A NEW BIT-LEVEL PIPELINED ARCHITECTURE FOR 2-D ALLPASS DIGITAL FILTERS", 1995 IEEE INTERNATIONAL SYMPOSIUM ON CIRCUITS AND SYSTEMS (ISCAS), SEATTLE, APR. 30 - MAY 3, 1995, vol. 1, 30 April 1995 (1995-04-30), INSTITUTE OF ELECTRICAL AND ELECTRONICS ENGINEERS, pages 724 - 727, XP000583315 *
K. MANIVANNAN ET AL.: "Minimal Multiplier Realization of 2-D All-Pass Digital Filters", IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS., vol. 35, no. 4, April 1988 (1988-04-01), NEW YORK US, pages 480 - 484, XP002104516 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1650869A1 (en) * 2004-10-20 2006-04-26 STMicroelectronics Pvt. Ltd A device for implementing a sum of products expression

Also Published As

Publication number Publication date
US20050193046A1 (en) 2005-09-01
EP1119909A1 (en) 2001-08-01
DE69821144D1 (en) 2004-02-19
DE69821144T2 (en) 2004-09-02
EP1119909B1 (en) 2004-01-14

Similar Documents

Publication Publication Date Title
EP0693236B1 (en) Method and arrangement in a transposed digital fir filter for multiplying a binary input signal with tap coefficients and a method for designing a transposed digital filter
JP4445132B2 (en) Digital filtering without multiplier
US6308190B1 (en) Low-power pulse-shaping digital filters
JPS6364100B2 (en)
Narendiran et al. An efficient modified distributed arithmetic architecture suitable for FIR filter
US7228325B2 (en) Bypassable adder
EP1119910B1 (en) Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output
EP1119909A1 (en) Area efficient realization of coefficient architecture for bit-serial fir, iir filters and combinational/sequential logic structure with zero latency clock output
EP0791242A1 (en) Improved digital filter
Abinaya et al. Heuristic Analysis of Multiplierless Desensitized Half-Band Decimation Filter for Wireless Applications
KR0140805B1 (en) Bit-serial operation unit
McNally et al. Optimized bit level architectures for IIR filtering
Ghanekar et al. A class of high-precision multiplier-free FIR filter realizations with periodically time-varying coefficients
Jones Efficient computation of time-varying and adaptive filters
Mehendale et al. Coefficient transformations for area-efficient implementation of multiplier-less FIR filters
Al-Haj An efficient configurable hardware implementation of fundamental multirate filter banks
Ghanekar et al. Multiplier-free IIR filter realization using periodically time-varying state-space structure. I. Structure and design
KR100369337B1 (en) Halfband Linear Phase Finite Impulse Response Filters
Padir Digital incremental recursive filtering using digital differential analyzers
Awasthi et al. Performance Analysis of Compensated CIC Filter in Efficient Computing Using Signed Digit Number System
JPH02108319A (en) Infinite impulse response type digital filter
JPH03211910A (en) digital filter
Francesconi et al. A novel interpolator architecture for/spl Sigma//spl Delta/DACs
JPH03159414A (en) Digital filter
JPH0381326B2 (en)

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A1

Designated state(s): JP SG US

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): AT BE CH CY DE DK ES FI FR GB GR IE IT LU MC NL PT SE

121 Ep: the epo has been informed by wipo that ep was designated in this application
DFPE Request for preliminary examination filed prior to expiration of 19th month from priority date (pct application filed before 20040101)
WWE Wipo information: entry into national phase

Ref document number: 1998950601

Country of ref document: EP

WWE Wipo information: entry into national phase

Ref document number: 09807498

Country of ref document: US

WWP Wipo information: published in national office

Ref document number: 1998950601

Country of ref document: EP

WWG Wipo information: grant in national office

Ref document number: 1998950601

Country of ref document: EP