On Feb 13, 2006, at 11:31 AM, kernels_nz wrote:
> 2. if you had :
>
> mydata[0] = PIND;
> mydata[1] = PIND;
> mydata[2] = PIND;
> .
> .
> .
> mydata[255] = PIND;
>
> all written out the hard way like that, it would be slightly faster,
> because the "for" loop condition would not be checked every time, it
> would of course occupy much more code space, and be considered
> terrible programming.
Yes, but there's still a middle way, which is partial unrolling:
#define ATONCE 8
uint8_t mydata[256];
void myloop(void)
{
uint8_t *p = mydata;
for (uint8_t counter = sizeof(mydata) / ATONCE; counter--; p +=
ATONCE)
{
p[0] = PIND;
p[1] = PIND;
p[2] = PIND;
p[3] = PIND;
p[4] = PIND;
p[5] = PIND;
p[6] = PIND;
p[7] = PIND;
}
}
which results in (with avr-gcc 4.0 and -O2):
9:test3.c **** void myloop(void)
10:test3.c **** {
74 .LM0:
75 /* prologue: frame size=0 */
76 /* prologue end (size=0) */
77 0000 E0E0 ldi r30,lo8(mydata)
78 0002 F0E0 ldi r31,hi8(mydata)
79 .L2:
80 .LBB2:
11:test3.c **** uint8_t *p = mydata;
12:test3.c **** for (uint8_t counter = sizeof(mydata) /
ATONCE; counter--; p += ATONCE)
13:test3.c **** {
14:test3.c **** p[0] = PIND;
82 .LM1:
83 0004 80B3 in r24,48-0x20
84 0006 8083 st Z,r24
15:test3.c **** p[1] = PIND;
86 .LM2:
87 0008 80B3 in r24,48-0x20
88 000a 8183 std Z+1,r24
16:test3.c **** p[2] = PIND;
90 .LM3:
91 000c 80B3 in r24,48-0x20
92 000e 8283 std Z+2,r24
... etc ...
21:test3.c **** p[7] = PIND;
110 .LM8:
111 0020 80B3 in r24,48-0x20
112 0022 8783 std Z+7,r24
114 .LM9:
115 0024 3896 adiw r30,8 ; 2
116 0026 80E0 ldi r24,hi8(mydata+256) ; 1
117 0028 E030 cpi r30,lo8(mydata+256) ; 1
118 002a F807 cpc r31,r24 ; 1
119 002c 59F7 brne .L2 ; 2
As a result, you amortize the cost of lines 115-119 (7 cycles) over 8
bytes of transfer.
For a cost/byte transferred of 3+7/8 cycles.
--
Ned Konz
ned@bike-nomad.comMessage
Re: [AVR-Chat] Re: Help needed:- what's the quickest way to store 256 bytes of data?
2006-02-14 by Ned Konz
Attachments
- No local attachments were found for this message.