Peter Zijlstra
1c34496e58
objtool: Remove instruction::list
...
Replace the instruction::list by allocating instructions in arrays of
256 entries and stringing them together by (amortized) find_insn().
This shrinks instruction by 16 bytes and brings it down to 128.
struct instruction {
- struct list_head list; /* 0 16 */
- struct hlist_node hash; /* 16 16 */
- struct list_head call_node; /* 32 16 */
- struct section * sec; /* 48 8 */
- long unsigned int offset; /* 56 8 */
- /* --- cacheline 1 boundary (64 bytes) --- */
- long unsigned int immediate; /* 64 8 */
- unsigned int len; /* 72 4 */
- u8 type; /* 76 1 */
-
- /* Bitfield combined with previous fields */
+ struct hlist_node hash; /* 0 16 */
+ struct list_head call_node; /* 16 16 */
+ struct section * sec; /* 32 8 */
+ long unsigned int offset; /* 40 8 */
+ long unsigned int immediate; /* 48 8 */
+ u8 len; /* 56 1 */
+ u8 prev_len; /* 57 1 */
+ u8 type; /* 58 1 */
+ s8 instr; /* 59 1 */
+ u32 idx:8; /* 60: 0 4 */
+ u32 dead_end:1; /* 60: 8 4 */
+ u32 ignore:1; /* 60: 9 4 */
+ u32 ignore_alts:1; /* 60:10 4 */
+ u32 hint:1; /* 60:11 4 */
+ u32 save:1; /* 60:12 4 */
+ u32 restore:1; /* 60:13 4 */
+ u32 retpoline_safe:1; /* 60:14 4 */
+ u32 noendbr:1; /* 60:15 4 */
+ u32 entry:1; /* 60:16 4 */
+ u32 visited:4; /* 60:17 4 */
+ u32 no_reloc:1; /* 60:21 4 */
- u16 dead_end:1; /* 76: 8 2 */
- u16 ignore:1; /* 76: 9 2 */
- u16 ignore_alts:1; /* 76:10 2 */
- u16 hint:1; /* 76:11 2 */
- u16 save:1; /* 76:12 2 */
- u16 restore:1; /* 76:13 2 */
- u16 retpoline_safe:1; /* 76:14 2 */
- u16 noendbr:1; /* 76:15 2 */
- u16 entry:1; /* 78: 0 2 */
- u16 visited:4; /* 78: 1 2 */
- u16 no_reloc:1; /* 78: 5 2 */
+ /* XXX 10 bits hole, try to pack */
- /* XXX 2 bits hole, try to pack */
- /* Bitfield combined with next fields */
-
- s8 instr; /* 79 1 */
- struct alt_group * alt_group; /* 80 8 */
- struct instruction * jump_dest; /* 88 8 */
- struct instruction * first_jump_src; /* 96 8 */
+ /* --- cacheline 1 boundary (64 bytes) --- */
+ struct alt_group * alt_group; /* 64 8 */
+ struct instruction * jump_dest; /* 72 8 */
+ struct instruction * first_jump_src; /* 80 8 */
union {
- struct symbol * _call_dest; /* 104 8 */
- struct reloc * _jump_table; /* 104 8 */
- }; /* 104 8 */
- struct alternative * alts; /* 112 8 */
- struct symbol * sym; /* 120 8 */
- /* --- cacheline 2 boundary (128 bytes) --- */
- struct stack_op * stack_ops; /* 128 8 */
- struct cfi_state * cfi; /* 136 8 */
+ struct symbol * _call_dest; /* 88 8 */
+ struct reloc * _jump_table; /* 88 8 */
+ }; /* 88 8 */
+ struct alternative * alts; /* 96 8 */
+ struct symbol * sym; /* 104 8 */
+ struct stack_op * stack_ops; /* 112 8 */
+ struct cfi_state * cfi; /* 120 8 */
- /* size: 144, cachelines: 3, members: 28 */
- /* sum members: 142 */
- /* sum bitfield members: 14 bits, bit holes: 1, sum bit holes: 2 bits */
- /* last cacheline: 16 bytes */
+ /* size: 128, cachelines: 2, members: 29 */
+ /* sum members: 124 */
+ /* sum bitfield members: 22 bits, bit holes: 1, sum bit holes: 10 bits */
};
pre: 5:38.18 real, 213.25 user, 124.90 sys, 23449040 mem
post: 5:03.34 real, 210.75 user, 88.80 sys, 20241232 mem
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:44 +01:00
Peter Zijlstra
6ea17e848a
x86: Fix FILL_RETURN_BUFFER
...
With overlapping alternative validation fixed, objtool promptly
complains:
vmlinux.o: warning: objtool: __switch_to_asm+0x2c: stack layout conflict in alternatives: .altinstr_replacement+0x47
.rela.altinstructions:
000000000000009c 0000000200000002 R_X86_64_PC32 0000000000000000 .text + 16dc
00000000000000a0 0000000600000002 R_X86_64_PC32 0000000000000000 .altinstr_replacement + 3a
00000000000000a8 0000000200000002 R_X86_64_PC32 0000000000000000 .text + 16dc
00000000000000ac 0000000600000002 R_X86_64_PC32 0000000000000000 .altinstr_replacement + 66
.text:
00000000000016b0 <__switch_to_asm>:
16b0: f3 0f 1e fa endbr64
16b4: 55 push %rbp
16b5: 53 push %rbx
16b6: 41 54 push %r12
16b8: 41 55 push %r13
16ba: 41 56 push %r14
16bc: 41 57 push %r15
16be: 48 89 a7 18 0b 00 00 mov %rsp,0xb18(%rdi)
16c5: 48 8b a6 18 0b 00 00 mov 0xb18(%rsi),%rsp
16cc: 48 8b 9e 28 05 00 00 mov 0x528(%rsi),%rbx
16d3: 65 48 89 1c 25 00 00 00 00 mov %rbx,%gs:0x0 16d8: R_X86_64_32S fixed_percpu_data+0x28
16dc: eb 2a jmp 1708 <__switch_to_asm+0x58>
16de: 90 nop
16df: 90 nop
16e0: 90 nop
16e1: 90 nop
16e2: 90 nop
16e3: 90 nop
16e4: 90 nop
16e5: 90 nop
16e6: 90 nop
16e7: 90 nop
16e8: 90 nop
16e9: 90 nop
16ea: 90 nop
16eb: 90 nop
16ec: 90 nop
16ed: 90 nop
16ee: 90 nop
16ef: 90 nop
16f0: 90 nop
16f1: 90 nop
16f2: 90 nop
16f3: 90 nop
16f4: 90 nop
16f5: 90 nop
16f6: 90 nop
16f7: 90 nop
16f8: 90 nop
16f9: 90 nop
16fa: 90 nop
16fb: 90 nop
16fc: 90 nop
16fd: 90 nop
16fe: 90 nop
16ff: 90 nop
1700: 90 nop
1701: 90 nop
1702: 90 nop
1703: 90 nop
1704: 90 nop
1705: 90 nop
1706: 90 nop
1707: 90 nop
1708: 41 5f pop %r15
170a: 41 5e pop %r14
170c: 41 5d pop %r13
170e: 41 5c pop %r12
1710: 5b pop %rbx
1711: 5d pop %rbp
1712: e9 00 00 00 00 jmp 1717 <__switch_to_asm+0x67> 1713: R_X86_64_PLT32 __switch_to-0x4
.altinstr_replacement:
3a: 49 c7 c4 10 00 00 00 mov $0x10,%r12
41: e8 01 00 00 00 call 47 <.altinstr_replacement+0x47>
46: cc int3
47: e8 01 00 00 00 call 4d <.altinstr_replacement+0x4d>
4c: cc int3
4d: 48 83 c4 10 add $0x10,%rsp
51: 49 ff cc dec %r12
54: 75 eb jne 41 <.altinstr_replacement+0x41>
56: 0f ae e8 lfence
59: 65 48 c7 04 25 00 00 00 00 ff ff ff ff movq $0xffffffffffffffff,%gs:0x0 5e: R_X86_64_32S pcpu_hot+0x10
66: e8 01 00 00 00 call 6c <.altinstr_replacement+0x6c>
6b: cc int3
6c: 48 83 c4 08 add $0x8,%rsp
70: 0f ae e8 lfence
As can be seen from the two alternatives, when overlaid, the NOP after
the shorter (starting at 66) coinsides with the call at 47, leading to
conflicting CFI state for that instruction.
By offsetting the shorter alternative by 2 bytes, this alignment is
undone.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:37 +01:00
Peter Zijlstra
a706bb08c8
objtool: Fix overlapping alternatives
...
Things like ALTERNATIVE_{2,3}() generate multiple alternatives on the
same place, objtool would override the first orig_alt_group with the
second (or third), failing to check the CFI among all the different
variants.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:33 +01:00
Peter Zijlstra
c6f5dc28fb
objtool: Union instruction::{call_dest,jump_table}
...
The instruction call_dest and jump_table members can never be used at
the same time, their usage depends on type.
struct instruction {
struct list_head list; /* 0 16 */
struct hlist_node hash; /* 16 16 */
struct list_head call_node; /* 32 16 */
struct section * sec; /* 48 8 */
long unsigned int offset; /* 56 8 */
/* --- cacheline 1 boundary (64 bytes) --- */
long unsigned int immediate; /* 64 8 */
unsigned int len; /* 72 4 */
u8 type; /* 76 1 */
/* Bitfield combined with previous fields */
u16 dead_end:1; /* 76: 8 2 */
u16 ignore:1; /* 76: 9 2 */
u16 ignore_alts:1; /* 76:10 2 */
u16 hint:1; /* 76:11 2 */
u16 save:1; /* 76:12 2 */
u16 restore:1; /* 76:13 2 */
u16 retpoline_safe:1; /* 76:14 2 */
u16 noendbr:1; /* 76:15 2 */
u16 entry:1; /* 78: 0 2 */
u16 visited:4; /* 78: 1 2 */
u16 no_reloc:1; /* 78: 5 2 */
/* XXX 2 bits hole, try to pack */
/* Bitfield combined with next fields */
s8 instr; /* 79 1 */
struct alt_group * alt_group; /* 80 8 */
- struct symbol * call_dest; /* 88 8 */
- struct instruction * jump_dest; /* 96 8 */
- struct instruction * first_jump_src; /* 104 8 */
- struct reloc * jump_table; /* 112 8 */
- struct alternative * alts; /* 120 8 */
+ struct instruction * jump_dest; /* 88 8 */
+ struct instruction * first_jump_src; /* 96 8 */
+ union {
+ struct symbol * _call_dest; /* 104 8 */
+ struct reloc * _jump_table; /* 104 8 */
+ }; /* 104 8 */
+ struct alternative * alts; /* 112 8 */
+ struct symbol * sym; /* 120 8 */
/* --- cacheline 2 boundary (128 bytes) --- */
- struct symbol * sym; /* 128 8 */
- struct stack_op * stack_ops; /* 136 8 */
- struct cfi_state * cfi; /* 144 8 */
+ struct stack_op * stack_ops; /* 128 8 */
+ struct cfi_state * cfi; /* 136 8 */
- /* size: 152, cachelines: 3, members: 29 */
- /* sum members: 150 */
+ /* size: 144, cachelines: 3, members: 28 */
+ /* sum members: 142 */
/* sum bitfield members: 14 bits, bit holes: 1, sum bit holes: 2 bits */
- /* last cacheline: 24 bytes */
+ /* last cacheline: 16 bytes */
};
pre: 5:39.35 real, 215.58 user, 123.69 sys, 23448736 mem
post: 5:38.18 real, 213.25 user, 124.90 sys, 23449040 mem
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:27 +01:00
Peter Zijlstra
0932dbe1f5
objtool: Remove instruction::reloc
...
Instead of caching the reloc for each instruction, only keep a
negative cache of not having a reloc (by far the most common case).
struct instruction {
struct list_head list; /* 0 16 */
struct hlist_node hash; /* 16 16 */
struct list_head call_node; /* 32 16 */
struct section * sec; /* 48 8 */
long unsigned int offset; /* 56 8 */
/* --- cacheline 1 boundary (64 bytes) --- */
long unsigned int immediate; /* 64 8 */
unsigned int len; /* 72 4 */
u8 type; /* 76 1 */
/* Bitfield combined with previous fields */
u16 dead_end:1; /* 76: 8 2 */
u16 ignore:1; /* 76: 9 2 */
u16 ignore_alts:1; /* 76:10 2 */
u16 hint:1; /* 76:11 2 */
u16 save:1; /* 76:12 2 */
u16 restore:1; /* 76:13 2 */
u16 retpoline_safe:1; /* 76:14 2 */
u16 noendbr:1; /* 76:15 2 */
u16 entry:1; /* 78: 0 2 */
u16 visited:4; /* 78: 1 2 */
+ u16 no_reloc:1; /* 78: 5 2 */
- /* XXX 3 bits hole, try to pack */
+ /* XXX 2 bits hole, try to pack */
/* Bitfield combined with next fields */
s8 instr; /* 79 1 */
struct alt_group * alt_group; /* 80 8 */
struct symbol * call_dest; /* 88 8 */
struct instruction * jump_dest; /* 96 8 */
struct instruction * first_jump_src; /* 104 8 */
struct reloc * jump_table; /* 112 8 */
- struct reloc * reloc; /* 120 8 */
+ struct alternative * alts; /* 120 8 */
/* --- cacheline 2 boundary (128 bytes) --- */
- struct alternative * alts; /* 128 8 */
- struct symbol * sym; /* 136 8 */
- struct stack_op * stack_ops; /* 144 8 */
- struct cfi_state * cfi; /* 152 8 */
+ struct symbol * sym; /* 128 8 */
+ struct stack_op * stack_ops; /* 136 8 */
+ struct cfi_state * cfi; /* 144 8 */
- /* size: 160, cachelines: 3, members: 29 */
- /* sum members: 158 */
- /* sum bitfield members: 13 bits, bit holes: 1, sum bit holes: 3 bits */
- /* last cacheline: 32 bytes */
+ /* size: 152, cachelines: 3, members: 29 */
+ /* sum members: 150 */
+ /* sum bitfield members: 14 bits, bit holes: 1, sum bit holes: 2 bits */
+ /* last cacheline: 24 bytes */
};
pre: 5:48.89 real, 220.96 user, 127.55 sys, 24834672 mem
post: 5:39.35 real, 215.58 user, 123.69 sys, 23448736 mem
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:17 +01:00
Peter Zijlstra
8b2de41215
objtool: Shrink instruction::{type,visited}
...
Since we don't have that many types in enum insn_type, force it into a
u8 and re-arrange member to get rid of the holes, saves another 8
bytes.
struct instruction {
struct list_head list; /* 0 16 */
struct hlist_node hash; /* 16 16 */
struct list_head call_node; /* 32 16 */
struct section * sec; /* 48 8 */
long unsigned int offset; /* 56 8 */
/* --- cacheline 1 boundary (64 bytes) --- */
- unsigned int len; /* 64 4 */
- enum insn_type type; /* 68 4 */
- long unsigned int immediate; /* 72 8 */
- u16 dead_end:1; /* 80: 0 2 */
- u16 ignore:1; /* 80: 1 2 */
- u16 ignore_alts:1; /* 80: 2 2 */
- u16 hint:1; /* 80: 3 2 */
- u16 save:1; /* 80: 4 2 */
- u16 restore:1; /* 80: 5 2 */
- u16 retpoline_safe:1; /* 80: 6 2 */
- u16 noendbr:1; /* 80: 7 2 */
- u16 entry:1; /* 80: 8 2 */
+ long unsigned int immediate; /* 64 8 */
+ unsigned int len; /* 72 4 */
+ u8 type; /* 76 1 */
- /* XXX 7 bits hole, try to pack */
+ /* Bitfield combined with previous fields */
- s8 instr; /* 82 1 */
- u8 visited; /* 83 1 */
+ u16 dead_end:1; /* 76: 8 2 */
+ u16 ignore:1; /* 76: 9 2 */
+ u16 ignore_alts:1; /* 76:10 2 */
+ u16 hint:1; /* 76:11 2 */
+ u16 save:1; /* 76:12 2 */
+ u16 restore:1; /* 76:13 2 */
+ u16 retpoline_safe:1; /* 76:14 2 */
+ u16 noendbr:1; /* 76:15 2 */
+ u16 entry:1; /* 78: 0 2 */
+ u16 visited:4; /* 78: 1 2 */
- /* XXX 4 bytes hole, try to pack */
+ /* XXX 3 bits hole, try to pack */
+ /* Bitfield combined with next fields */
- struct alt_group * alt_group; /* 88 8 */
- struct symbol * call_dest; /* 96 8 */
- struct instruction * jump_dest; /* 104 8 */
- struct instruction * first_jump_src; /* 112 8 */
- struct reloc * jump_table; /* 120 8 */
+ s8 instr; /* 79 1 */
+ struct alt_group * alt_group; /* 80 8 */
+ struct symbol * call_dest; /* 88 8 */
+ struct instruction * jump_dest; /* 96 8 */
+ struct instruction * first_jump_src; /* 104 8 */
+ struct reloc * jump_table; /* 112 8 */
+ struct reloc * reloc; /* 120 8 */
/* --- cacheline 2 boundary (128 bytes) --- */
- struct reloc * reloc; /* 128 8 */
- struct alternative * alts; /* 136 8 */
- struct symbol * sym; /* 144 8 */
- struct stack_op * stack_ops; /* 152 8 */
- struct cfi_state * cfi; /* 160 8 */
+ struct alternative * alts; /* 128 8 */
+ struct symbol * sym; /* 136 8 */
+ struct stack_op * stack_ops; /* 144 8 */
+ struct cfi_state * cfi; /* 152 8 */
- /* size: 168, cachelines: 3, members: 29 */
- /* sum members: 162, holes: 1, sum holes: 4 */
- /* sum bitfield members: 9 bits, bit holes: 1, sum bit holes: 7 bits */
- /* last cacheline: 40 bytes */
+ /* size: 160, cachelines: 3, members: 29 */
+ /* sum members: 158 */
+ /* sum bitfield members: 13 bits, bit holes: 1, sum bit holes: 3 bits */
+ /* last cacheline: 32 bytes */
};
pre: 5:48.86 real, 220.30 user, 128.34 sys, 24834672 mem
post: 5:48.89 real, 220.96 user, 127.55 sys, 24834672 mem
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:12 +01:00
Peter Zijlstra
d540665461
objtool: Make instruction::alts a single-linked list
...
struct instruction {
struct list_head list; /* 0 16 */
struct hlist_node hash; /* 16 16 */
struct list_head call_node; /* 32 16 */
struct section * sec; /* 48 8 */
long unsigned int offset; /* 56 8 */
/* --- cacheline 1 boundary (64 bytes) --- */
unsigned int len; /* 64 4 */
enum insn_type type; /* 68 4 */
long unsigned int immediate; /* 72 8 */
u16 dead_end:1; /* 80: 0 2 */
u16 ignore:1; /* 80: 1 2 */
u16 ignore_alts:1; /* 80: 2 2 */
u16 hint:1; /* 80: 3 2 */
u16 save:1; /* 80: 4 2 */
u16 restore:1; /* 80: 5 2 */
u16 retpoline_safe:1; /* 80: 6 2 */
u16 noendbr:1; /* 80: 7 2 */
u16 entry:1; /* 80: 8 2 */
/* XXX 7 bits hole, try to pack */
s8 instr; /* 82 1 */
u8 visited; /* 83 1 */
/* XXX 4 bytes hole, try to pack */
struct alt_group * alt_group; /* 88 8 */
struct symbol * call_dest; /* 96 8 */
struct instruction * jump_dest; /* 104 8 */
struct instruction * first_jump_src; /* 112 8 */
struct reloc * jump_table; /* 120 8 */
/* --- cacheline 2 boundary (128 bytes) --- */
struct reloc * reloc; /* 128 8 */
- struct list_head alts; /* 136 16 */
- struct symbol * sym; /* 152 8 */
- struct stack_op * stack_ops; /* 160 8 */
- struct cfi_state * cfi; /* 168 8 */
+ struct alternative * alts; /* 136 8 */
+ struct symbol * sym; /* 144 8 */
+ struct stack_op * stack_ops; /* 152 8 */
+ struct cfi_state * cfi; /* 160 8 */
- /* size: 176, cachelines: 3, members: 29 */
- /* sum members: 170, holes: 1, sum holes: 4 */
+ /* size: 168, cachelines: 3, members: 29 */
+ /* sum members: 162, holes: 1, sum holes: 4 */
/* sum bitfield members: 9 bits, bit holes: 1, sum bit holes: 7 bits */
- /* last cacheline: 48 bytes */
+ /* last cacheline: 40 bytes */
};
pre: 5:58.50 real, 229.64 user, 128.65 sys, 26221520 mem
post: 5:48.86 real, 220.30 user, 128.34 sys, 24834672 mem
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:21:06 +01:00
Peter Zijlstra
3ee88df1b0
objtool: Make instruction::stack_ops a single-linked list
...
struct instruction {
struct list_head list; /* 0 16 */
struct hlist_node hash; /* 16 16 */
struct list_head call_node; /* 32 16 */
struct section * sec; /* 48 8 */
long unsigned int offset; /* 56 8 */
/* --- cacheline 1 boundary (64 bytes) --- */
unsigned int len; /* 64 4 */
enum insn_type type; /* 68 4 */
long unsigned int immediate; /* 72 8 */
u16 dead_end:1; /* 80: 0 2 */
u16 ignore:1; /* 80: 1 2 */
u16 ignore_alts:1; /* 80: 2 2 */
u16 hint:1; /* 80: 3 2 */
u16 save:1; /* 80: 4 2 */
u16 restore:1; /* 80: 5 2 */
u16 retpoline_safe:1; /* 80: 6 2 */
u16 noendbr:1; /* 80: 7 2 */
u16 entry:1; /* 80: 8 2 */
/* XXX 7 bits hole, try to pack */
s8 instr; /* 82 1 */
u8 visited; /* 83 1 */
/* XXX 4 bytes hole, try to pack */
struct alt_group * alt_group; /* 88 8 */
struct symbol * call_dest; /* 96 8 */
struct instruction * jump_dest; /* 104 8 */
struct instruction * first_jump_src; /* 112 8 */
struct reloc * jump_table; /* 120 8 */
/* --- cacheline 2 boundary (128 bytes) --- */
struct reloc * reloc; /* 128 8 */
struct list_head alts; /* 136 16 */
struct symbol * sym; /* 152 8 */
- struct list_head stack_ops; /* 160 16 */
- struct cfi_state * cfi; /* 176 8 */
+ struct stack_op * stack_ops; /* 160 8 */
+ struct cfi_state * cfi; /* 168 8 */
- /* size: 184, cachelines: 3, members: 29 */
- /* sum members: 178, holes: 1, sum holes: 4 */
+ /* size: 176, cachelines: 3, members: 29 */
+ /* sum members: 170, holes: 1, sum holes: 4 */
/* sum bitfield members: 9 bits, bit holes: 1, sum bit holes: 7 bits */
- /* last cacheline: 56 bytes */
+ /* last cacheline: 48 bytes */
};
pre: 5:58.22 real, 226.69 user, 131.22 sys, 26221520 mem
post: 5:58.50 real, 229.64 user, 128.65 sys, 26221520 mem
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:20:59 +01:00
Peter Zijlstra
20a554638d
objtool: Change arch_decode_instruction() signature
...
In preparation to changing struct instruction around a bit, avoid
passing it's members by pointer and instead pass the whole thing.
A cleanup in it's own right too.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Josh Poimboeuf <[email protected] >
Tested-by: Nathan Chancellor <[email protected] > # build only
Tested-by: Thomas Weißschuh <[email protected] > # compile and run
Link: https://lore.kernel.org/r/[email protected]
2023-02-23 09:20:50 +01:00
Peter Zijlstra
eedeb787eb
freezer,umh: Fix call_usermode_helper_exec() vs SIGKILL
...
Tetsuo-San noted that commit f5d39b0208 ("freezer,sched: Rewrite
core freezer logic") broke call_usermodehelper_exec() for the KILLABLE
case.
Specifically it was missed that the second, unconditional,
wait_for_completion() was not optional and ensures the on-stack
completion is unused before going out-of-scope.
Fixes: f5d39b0208 ("freezer,sched: Rewrite core freezer logic")
Reported-by: [email protected]
Reported-by: Tetsuo Handa <[email protected] >
Debugged-by: Tetsuo Handa <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2023-02-13 16:36:14 +01:00
Peter Zijlstra
443ed4c302
objtool: mem*() are not uaccess safe
...
For mysterious raisins I listed the new __asan_mem*() functions as
being uaccess safe, this is giving objtool fails on KASAN builds
because these functions call out to the actual __mem*() functions
which are not marked uaccess safe.
Removing it doesn't make the robots unhappy.
Fixes: 69d4c0d321 ("entry, kasan, x86: Disallow overriding mem*() functions")
Reported-by: "Paul E. McKenney" <[email protected] >
Bisected-by: Josh Poimboeuf <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20230126182302.GA687063@paulmck-ThinkPad-P17-Gen-1
2023-02-11 11:18:08 +01:00
Peter Zijlstra
923510c88d
x86/static_call: Add support for Jcc tail-calls
...
Clang likes to create conditional tail calls like:
0000000000000350 <amd_pmu_add_event>:
350: 0f 1f 44 00 00 nopl 0x0(%rax,%rax,1) 351: R_X86_64_NONE __fentry__-0x4
355: 48 83 bf 20 01 00 00 00 cmpq $0x0,0x120(%rdi)
35d: 0f 85 00 00 00 00 jne 363 <amd_pmu_add_event+0x13> 35f: R_X86_64_PLT32 __SCT__amd_pmu_branch_add-0x4
363: e9 00 00 00 00 jmp 368 <amd_pmu_add_event+0x18> 364: R_X86_64_PLT32 __x86_return_thunk-0x4
Where 0x35d is a static call site that's turned into a conditional
tail-call using the Jcc class of instructions.
Teach the in-line static call text patching about this.
Notably, since there is no conditional-ret, in that case patch the Jcc
to point at an empty stub function that does the ret -- or the return
thunk when needed.
Reported-by: "Erhard F." <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Reviewed-by: Masami Hiramatsu (Google) <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:05:31 +01:00
Peter Zijlstra
ac0ee0a956
x86/alternatives: Teach text_poke_bp() to patch Jcc.d32 instructions
...
In order to re-write Jcc.d32 instructions text_poke_bp() needs to be
taught about them.
The biggest hurdle is that the whole machinery is currently made for 5
byte instructions and extending this would grow struct text_poke_loc
which is currently a nice 16 bytes and used in an array.
However, since text_poke_loc contains a full copy of the (s32)
displacement, it is possible to map the Jcc.d32 2 byte opcodes to
Jcc.d8 1 byte opcode for the int3 emulation.
This then leaves the replacement bytes; fudge that by only storing the
last 5 bytes and adding the rule that 'length == 6' instruction will
be prefixed with a 0x0f byte.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Reviewed-by: Masami Hiramatsu (Google) <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:05:31 +01:00
Peter Zijlstra
db7adcfd1c
x86/alternatives: Introduce int3_emulate_jcc()
...
Move the kprobe Jcc emulation into int3_emulate_jcc() so it can be
used by more code -- specifically static_call() will need this.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Reviewed-by: Masami Hiramatsu (Google) <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:05:30 +01:00
Peter Zijlstra
4d627628d7
cpuidle: Fix poll_idle() noinstr annotation
...
The instrumentation_begin()/end() annotations in poll_idle() were
complete nonsense. Specifically they caused tracing to happen in the
middle of noinstr code, resulting in RCU splats.
Now that local_clock() is noinstr, mark up the rest and let it rip.
Fixes: 00717eb8c9 ("cpuidle: Annotate poll_idle()")
Reported-by: kernel test robot <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/oe-lkp/[email protected]
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:47 +01:00
Peter Zijlstra
776f22913b
sched/clock: Make local_clock() noinstr
...
With sched_clock() noinstr, provide a noinstr implementation of
local_clock().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:47 +01:00
Peter Zijlstra
8739c68115
sched/clock/x86: Mark sched_clock() noinstr
...
In order to use sched_clock() from noinstr code, mark it and all it's
implenentations noinstr.
The whole pvclock thing (used by KVM/Xen) is a bit of a pain,
since it calls out to watchdogs, create a
pvclock_clocksource_read_nowd() variant doesn't do that and can be
noinstr.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:47 +01:00
Peter Zijlstra
7aab7aa4b4
x86/atomics: Always inline arch_atomic64*()
...
As already done for regular arch_atomic*(), always inline arch_atomic64*().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:46 +01:00
Peter Zijlstra
3017ba4b83
cpuidle: tracing, preempt: Squash _rcuidle tracing
...
Extend/fix commit:
9aedeaed6f ("tracing, hardirq: No moar _rcuidle() tracing")
... to also cover trace_preempt_{on,off}() which were mysteriously
untouched.
Fixes: 9aedeaed6f ("tracing, hardirq: No moar _rcuidle() tracing")
Reported-by: Mark Rutland <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Mark Rutland <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:46 +01:00
Peter Zijlstra
d099dbfd33
cpuidle: tracing: Warn about !rcu_is_watching()
...
When using noinstr, WARN when tracing hits when RCU is disabled.
Suggested-by: Steven Rostedt (Google) <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:46 +01:00
Peter Zijlstra
5a5d7e9bad
cpuidle: lib/bug: Disable rcu_is_watching() during WARN/BUG
...
In order to avoid WARN/BUG from generating nested or even recursive
warnings, force rcu_is_watching() true during
WARN/lockdep_rcu_suspicious().
Notably things like unwinding the stack can trigger rcu_dereference()
warnings, which then triggers more unwinding which then triggers more
warnings etc..
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-31 15:01:45 +01:00
Peter Zijlstra
19235e4727
cpuidle, arm64: Fix the ARM64 cpuidle logic
...
The recent cpuidle changes started triggering RCU splats on
Juno development boards:
| =============================
| WARNING: suspicious RCU usage
| -----------------------------
| include/trace/events/ipi.h:19 suspicious rcu_dereference_check() usage!
Fix cpuidle on ARM64:
- ... by introducing a new 'is_rcu' flag to the cpuidle helpers & make
ARM64 use it, as ARM64 wants to keep RCU active longer and wants to
do the ct_cpuidle_enter()/exit() dance itself.
- Also update the PSCI driver accordingly.
- This also removes the last known RCU_NONIDLE() user as a bonus.
Reported-by: Mark Rutland <[email protected] >
Signed-off-by: Peter Zijlstra <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Sudeep Holla <[email protected] >
Tested-by: Mark Rutland <[email protected] >
Reviewed-by: Mark Rutland <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
--
2023-01-18 12:27:17 +01:00
Peter Zijlstra
0e26e1de00
context_tracking: Fix noinstr vs KASAN
...
Low level noinstr context-tracking code is calling out to instrumented
code on KASAN:
vmlinux.o: warning: objtool: __ct_user_enter+0x72: call to __kasan_check_write() leaves .noinstr.text section
vmlinux.o: warning: objtool: __ct_user_exit+0x47: call to __kasan_check_write() leaves .noinstr.text section
Use even lower level atomic methods to avoid the instrumentation.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:18 +01:00
Peter Zijlstra
0e985e9d22
cpuidle: Add comments about noinstr/__cpuidle usage
...
Add a few words on noinstr / __cpuidle usage.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:18 +01:00
Peter Zijlstra
26388a7c35
cpuidle,arch: Mark all regular cpuidle_state:: Enter methods __cpuidle
...
For all cpuidle drivers that do not use CPUIDLE_FLAG_RCU_IDLE (iow,
the simple ones) make sure all the functions are marked __cpuidle.
( due to lack of noinstr validation on these platforms it is entirely
possible this isn't complete )
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:18 +01:00
Peter Zijlstra
69e26b4f43
cpuidle, arch: Mark all ct_cpuidle_enter() callers __cpuidle
...
For all cpuidle drivers that use CPUIDLE_FLAG_RCU_IDLE, ensure that
all functions that call ct_cpuidle_enter() are marked __cpuidle.
( due to lack of noinstr validation on these platforms it is entirely
possible this isn't complete )
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
17cc2b5525
cpuidle: Ensure ct_cpuidle_enter() is always called from noinstr/__cpuidle
...
Tracing (kprobes included) and other compiler instrumentation relies
on a normal kernel runtime. Therefore all functions that disable RCU
should be noinstr, as should all functions that are called while RCU
is disabled.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
1c38b0615f
arm64, riscv, perf: Remove RCU_NONIDLE() usage
...
The PM notifiers should no longer be ran with RCU disabled (per the
previous patches), as such this hack is no longer required either.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
f176d4ccb3
sched/core: Always inline __this_cpu_preempt_check()
...
Quite a few unnecessary instrumentation calls are generated via the
no-op __this_cpu_preempt_check() call, if it gets uninlined by the
compiler:
vmlinux.o: warning: objtool: in_entry_stack+0x9: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: default_do_nmi+0x10: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: fpu_idle_fpregs+0x41: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: kvm_read_and_reset_apf_flags+0x1: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: lockdep_hardirqs_on+0xb0: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: lockdep_hardirqs_off+0xae: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: irqentry_nmi_enter+0x69: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: irqentry_nmi_exit+0x32: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_processor_ffh_cstate_enter+0x9: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_idle_enter+0x43: call to __this_cpu_preempt_check() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_idle_enter_s2idle+0x45: call to __this_cpu_preempt_check() leaves .noinstr.text section
Mark it __always_inline.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
69d4c0d321
entry, kasan, x86: Disallow overriding mem*() functions
...
KASAN cannot just hijack the mem*() functions, it needs to emit
__asan_mem*() variants if it wants instrumentation (other sanitizers
already do this).
vmlinux.o: warning: objtool: sync_regs+0x24: call to memcpy() leaves .noinstr.text section
vmlinux.o: warning: objtool: vc_switch_off_ist+0xbe: call to memcpy() leaves .noinstr.text section
vmlinux.o: warning: objtool: fixup_bad_iret+0x36: call to memset() leaves .noinstr.text section
vmlinux.o: warning: objtool: __sev_get_ghcb+0xa0: call to memcpy() leaves .noinstr.text section
vmlinux.o: warning: objtool: __sev_put_ghcb+0x35: call to memcpy() leaves .noinstr.text section
Remove the weak aliases to ensure nobody hijacks these functions and
add them to the noinstr section.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
365bd03ff6
intel_idle: Add force_irq_on module param
...
For testing purposes.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
f18b0d7ee8
ubsan: Fix objtool UACCESS warns
...
clang-14 allyesconfig gives:
vmlinux.o: warning: objtool: emulator_cmpxchg_emulated+0x705: call to __ubsan_handle_load_invalid_value() with UACCESS enabled
vmlinux.o: warning: objtool: paging64_update_accessed_dirty_bits+0x39e: call to __ubsan_handle_load_invalid_value() with UACCESS enabled
vmlinux.o: warning: objtool: paging32_update_accessed_dirty_bits+0x390: call to __ubsan_handle_load_invalid_value() with UACCESS enabled
vmlinux.o: warning: objtool: ept_update_accessed_dirty_bits+0x43f: call to __ubsan_handle_load_invalid_value() with UACCESS enabled
Add the required eflags save/restore and whitelist the thing.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
ca502fc6d9
cpuidle, clk: Remove trace_.*_rcuidle()
...
OMAP was the one and only user.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Acked-by: Stephen Boyd <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
db8f50861d
cpuidle, ARM: OMAP2+: powerdomain: Remove trace_.*_rcuidle()
...
OMAP was the one and only user.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
259c95afac
arm, OMAP2: Use WFI for omap2_pm_idle()
...
arch_cpu_idle() is a very simple idle interface and exposes only a
single idle state and is expected to not require RCU and not do any
tracing/instrumentation.
As such, omap2_pm_idle() is not a valid implementation. Replace it
with a simple (shallow) omap2_do_wfi() call.
Omap2 doesn't have a cpuidle driver; but adding one would be the
recourse to (re)gain the other idle states.
Suggested-by: Tony Lindgren <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:17 +01:00
Peter Zijlstra
8c0956aa76
cpuidle, OMAP3: Push RCU-idle into omap_sram_idle()
...
OMAP3 uses full SoC suspend modes as idle states, as such it needs the
whole power-domain and clock-domain code from the idle path.
All that code is not suitable to run with RCU disabled, as such push
RCU-idle deeper still.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Tony Lindgren <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
40dbea961a
cpuidle, OMAP3: Use WFI for omap3_pm_idle()
...
arch_cpu_idle() is a very simple idle interface and exposes only a
single idle state and is expected to not require RCU and not do any
tracing/instrumentation.
As such, omap_sram_idle() is not a valid implementation. Replace it
with the simple (shallow) omap3_do_wfi() call. Leaving the more
complicated idle states for the cpuidle driver.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Tony Lindgren <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
9aedeaed6f
tracing, hardirq: No moar _rcuidle() tracing
...
Robot reported that trace_hardirqs_{on,off}() tickle the forbidden
_rcuidle() tracepoint through local_irq_{en,dis}able().
For 'sane' configs, these calls will only happen with RCU enabled and
as such can use the regular tracepoint. This also means it's possible
to trace them from NMI context again.
Reported-by: kernel test robot <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
408b961146
tracing: WARN on rcuidle
...
ARCH_WANTS_NO_INSTR (a superset of CONFIG_GENERIC_ENTRY) disallows any
and all tracing when RCU isn't enabled.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
dc7305606d
tracing: Remove trace_hardirqs_{on,off}_caller()
...
Per commit 56e62a7370 ("s390: convert to generic entry") the last
and only callers of trace_hardirqs_{on,off}_caller() went away, clean
up.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
6a123d6ae6
cpuidle, ACPI: Make noinstr clean
...
objtool found cases where ACPI methods called out into instrumentation code:
vmlinux.o: warning: objtool: io_idle+0xc: call to __inb.isra.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_idle_enter+0xfe: call to num_online_cpus() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_idle_enter+0x115: call to acpi_idle_fallback_to_c1.isra.0() leaves .noinstr.text section
Fix this by: marking the IO in/out, acpi_idle_fallback_to_c1() and
num_online_cpus() methods as __always_inline.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
10fdb38cee
cpuidle, nospec: Make mds_idle_clear_cpu_buffers() noinstr clean
...
objtool found that the mds_idle_clear_cpu_buffers() method got
uninlined by the compiler where it called out into instrumentation:
vmlinux.o: warning: objtool: mwait_idle+0x47: call to mds_idle_clear_cpu_buffers() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_processor_ffh_cstate_enter+0xa2: call to mds_idle_clear_cpu_buffers() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle+0x91: call to mds_idle_clear_cpu_buffers() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_s2idle+0x8c: call to mds_idle_clear_cpu_buffers() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_irq+0xaa: call to mds_idle_clear_cpu_buffers() leaves .noinstr.text section
Solve this by marking mds_idle_clear_cpu_buffers() as __always_inline.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
10a099405f
cpuidle, xenpv: Make more PARAVIRT_XXL noinstr clean
...
objtool found a few cases where this code called out into instrumented
code:
vmlinux.o: warning: objtool: acpi_idle_enter_s2idle+0xde: call to wbinvd() leaves .noinstr.text section
vmlinux.o: warning: objtool: default_idle+0x4: call to arch_safe_halt() leaves .noinstr.text section
vmlinux.o: warning: objtool: xen_safe_halt+0xa: call to HYPERVISOR_sched_op.constprop.0() leaves .noinstr.text section
Solve this by:
- marking arch_safe_halt(), wbinvd(), native_wbinvd() and
HYPERVISOR_sched_op() as __always_inline().
- Explicitly uninlining xen_safe_halt() and pv_native_wbinvd() [they were
already uninlined by the compiler on use as function pointers] and
annotating them as 'noinstr'.
- Annotating pv_native_safe_halt() as 'noinstr'.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Srivatsa S. Bhat (VMware) <[email protected] >
Reviewed-by: Juergen Gross <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
c3982c1a36
cpuidle, tdx: Make TDX code noinstr clean
...
objtool found a few cases where this code called out into instrumented
code:
vmlinux.o: warning: objtool: __halt+0x2c: call to hcall_func.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: __halt+0x3f: call to __tdx_hypercall() leaves .noinstr.text section
vmlinux.o: warning: objtool: __tdx_hypercall+0x66: call to __tdx_hypercall_failed() leaves .noinstr.text section
Fix it by:
- moving TDX tdcall assembly methods into .noinstr.text (they are already noistr-clean)
- marking __tdx_hypercall_failed() as 'noinstr'
- annotating hcall_func() as __always_inline
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
2ec8efe64e
cpuidle, mwait: Make the mwait code noinstr clean
...
objtool found a few cases where this code called out into instrumented
code:
vmlinux.o: warning: objtool: intel_idle_s2idle+0x6e: call to __monitor.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_irq+0x8c: call to __monitor.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle+0x73: call to __monitor.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: mwait_idle+0x88: call to clflush() leaves .noinstr.text section
Fix it by marking the affected methods as __always_inline.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
e4df1511e1
cpuidle, sched: Remove instrumentation from TIF_{POLLING_NRFLAG,NEED_RESCHED}
...
objtool pointed out that various idle-TIF management methods
have instrumentation:
vmlinux.o: warning: objtool: mwait_idle+0x5: call to current_set_polling_and_test() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_processor_ffh_cstate_enter+0xc5: call to current_set_polling_and_test() leaves .noinstr.text section
vmlinux.o: warning: objtool: cpu_idle_poll.isra.0+0x73: call to test_ti_thread_flag() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle+0xbc: call to current_set_polling_and_test() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_irq+0xea: call to current_set_polling_and_test() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_s2idle+0xb4: call to current_set_polling_and_test() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle+0xa6: call to current_clr_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_irq+0xbf: call to current_clr_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_s2idle+0xa1: call to current_clr_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: mwait_idle+0xe: call to __current_set_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_processor_ffh_cstate_enter+0xc5: call to __current_set_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: cpu_idle_poll.isra.0+0x73: call to test_ti_thread_flag() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle+0xbc: call to __current_set_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_irq+0xea: call to __current_set_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_s2idle+0xb4: call to __current_set_polling() leaves .noinstr.text section
vmlinux.o: warning: objtool: cpu_idle_poll.isra.0+0x73: call to test_ti_thread_flag() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_s2idle+0x73: call to test_ti_thread_flag.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_irq+0x91: call to test_ti_thread_flag.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle+0x78: call to test_ti_thread_flag.constprop.0() leaves .noinstr.text section
vmlinux.o: warning: objtool: acpi_safe_halt+0xf: call to test_ti_thread_flag.constprop.0() leaves .noinstr.text section
Remove the instrumentation, because these methods are used in low-level
cpuidle code moving between states, that should not be instrumented.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
e3ee5e66f7
time/tick-broadcast: Remove RCU_NONIDLE() usage
...
No callers left that have already disabled RCU.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Mark Rutland <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
880970b56b
printk: Remove trace_.*_rcuidle() usage
...
The problem, per commit fc98c3c8c9 ("printk: use rcuidle console
tracepoint"), was printk usage from the cpuidle path where RCU was
already disabled.
Per the patches earlier in this series, this is no longer the case.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Sergey Senozhatsky <[email protected] >
Acked-by: Petr Mladek <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:16 +01:00
Peter Zijlstra
4a3182e6d6
arm64, smp: Remove trace_.*_rcuidle() usage
...
Ever since commit d3afc7f129 ("arm64: Allow IPIs to be handled as
normal interrupts") this function is called in regular IRQ context.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Mark Rutland <[email protected] >
Acked-by: Marc Zyngier <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
08a56e07cd
arm, smp: Remove trace_.*_rcuidle() usage
...
None of these functions should ever be ran with RCU disabled anymore.
Specifically, do_handle_IPI() is only called from handle_IPI() which
explicitly does irq_enter()/irq_exit() which ensures RCU is watching.
The problem with smp_cross_call() was, per commit description:
7c64cc0531 ("arm: Use _rcuidle for smp_cross_call() tracepoints")
... that cpuidle_enter_state_coupled() already had RCU disabled, but that's
long been fixed by commit:
1098582a0f ("sched,idle,rcu: Push rcu_idle deeper into the idle path")
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
e80a48bade
x86/tdx: Remove TDX_HCALL_ISSUE_STI
...
Now that arch_cpu_idle() is expected to return with IRQs disabled,
avoid the useless STI/CLI dance.
Per the specs this is supposed to work, but nobody has yet relied up
this behaviour so broken implementations are possible.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
89b3098703
arch/idle: Change arch_cpu_idle() behavior: always exit with IRQs disabled
...
Current arch_cpu_idle() is called with IRQs disabled, but will return
with IRQs enabled.
However, the very first thing the generic code does after calling
arch_cpu_idle() is raw_local_irq_disable(). This means that
architectures that can idle with IRQs disabled end up doing a
pointless 'enable-disable' dance.
Therefore, push this IRQ disabling into the idle function, meaning
that those architectures can avoid the pointless IRQ state flipping.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Gautham R. Shenoy <[email protected] >
Acked-by: Mark Rutland <[email protected] > [arm64]
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Guo Ren <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
9b461a6faa
cpuidle, intel_idle: Fix CPUIDLE_FLAG_IBRS
...
objtool to the rescue:
vmlinux.o: warning: objtool: intel_idle_ibrs+0x17: call to spec_ctrl_current() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_ibrs+0x27: call to wrmsrl.constprop.0() leaves .noinstr.text section
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
821ad23d0e
cpuidle, intel_idle: Fix CPUIDLE_FLAG_INIT_XSTATE
...
Fix instrumentation bugs objtool found:
vmlinux.o: warning: objtool: intel_idle_s2idle+0xd5: call to fpu_idle_fpregs() leaves .noinstr.text section
vmlinux.o: warning: objtool: intel_idle_xstate+0x11: call to fpu_idle_fpregs() leaves .noinstr.text section
vmlinux.o: warning: objtool: fpu_idle_fpregs+0x9: call to xfeatures_in_use() leaves .noinstr.text section
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
6d9c7f51b1
cpuidle, intel_idle: Fix CPUIDLE_FLAG_IRQ_ENABLE *again*
...
So objtool found this bug:
vmlinux.o: warning: objtool: intel_idle_irq+0x10c: call to trace_hardirqs_off() leaves .noinstr.text section
As per commit 32d4fd5751 ("cpuidle,intel_idle: Fix CPUIDLE_FLAG_IRQ_ENABLE"):
"must not have tracing in idle functions"
Clearly people can't read and tinker along until splat dissapears.
This straight up reverts commit d295ad34f2 ("intel_idle: Fix false
positive RCU splats due to incorrect hardirqs state").
It doesn't re-introduce the problem because preceding patches fixed it
properly.
Fixes: d295ad34f2 ("intel_idle: Fix false positive RCU splats due to incorrect hardirqs state")
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
2b5a0e425e
objtool/idle: Validate __cpuidle code as noinstr
...
Idle code is very like entry code in that RCU isn't available. As
such, add a little validation.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Geert Uytterhoeven <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
00717eb8c9
cpuidle: Annotate poll_idle()
...
The __cpuidle functions will become a noinstr class, as such they need
explicit annotations.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
8ce78470bf
acpi_idle: Remove tracing
...
All the idle routines are called with RCU disabled, as such there must
not be any tracing inside.
While there; clean-up the io-port idle thing.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
924aed1646
cpuidle, cpu_pm: Remove RCU fiddling from cpu_pm_{enter,exit}()
...
All callers should still have RCU enabled.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Ulf Hansson <[email protected] >
Acked-by: Mark Rutland <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
a01353cf18
cpuidle: Fix ct_idle_*() usage
...
The whole disable-RCU, enable-IRQS dance is very intricate since
changing IRQ state is traced, which depends on RCU.
Add two helpers for the cpuidle case that mirror the entry code:
ct_cpuidle_enter()
ct_cpuidle_exit()
And fix all the cases where the enter/exit dance was buggy.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:15 +01:00
Peter Zijlstra
0c5ffc3d7b
cpuidle, dt: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again before going idle is suboptimal.
Notably: this converts all dt_init_idle_driver() and
__CPU_PM_CPU_IDLE_ENTER() users for they are inextrably intertwined.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:48:00 +01:00
Peter Zijlstra
c3d42418dc
cpuidle, OMAP4: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again, some *four* times, before going idle is suboptimal.
Notably three times explicitly using RCU_NONIDLE() and once implicitly
through cpu_pm_*().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Reviewed-by: Tony Lindgren <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:47:49 +01:00
Peter Zijlstra
4ce40e9dbe
cpuidle, armada: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again before going idle is suboptimal.
Notably the cpu_pm_*() calls implicitly re-enable RCU for a bit.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:47:27 +01:00
Peter Zijlstra
4d1be9e745
cpuidle, OMAP3: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then teporarily enable it
again before going idle is suboptimal.
Notably the cpu_pm_*() calls implicitly re-enable RCU for a bit.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Reviewed-by: Tony Lindgren <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:47:19 +01:00
Peter Zijlstra
b3f46658ce
cpuidle, ARM/imx6: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again, at least twice, before going idle is suboptimal.
Notably both cpu_pm_enter() and cpu_cluster_pm_enter() implicity
re-enable RCU.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:46:40 +01:00
Peter Zijlstra
e038f7b802
cpuidle, psci: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again, at least twice, before going idle is suboptimal.
Notably once implicitly through the cpu_pm_*() calls and once
explicitly doing ct_irq_*_irqon().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Kajetan Puchalski <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Reviewed-by: Guo Ren <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:22 +01:00
Peter Zijlstra
5fca0d9f5d
cpuidle, tegra: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again, at least twice, before going idle is suboptimal.
Notably once implicitly through the cpu_pm_*() calls and once
explicitly doing RCU_NONIDLE().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:22 +01:00
Peter Zijlstra
8e9ab9e8da
cpuidle, riscv: Push RCU-idle into driver
...
Doing RCU-idle outside the driver, only to then temporarily enable it
again, at least twice, before going idle is suboptimal.
That is, once implicitly through the cpu_pm_*() calls and once
explicitly doing ct_irq_*_irqon().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Anup Patel <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:22 +01:00
Peter Zijlstra
bb7b112585
cpuidle: Move IRQ state validation
...
Make cpuidle_enter_state() consistent with the s2idle variant and
verify ->enter() always returns with interrupts disabled.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:22 +01:00
Peter Zijlstra
5e26aa9339
cpuidle/poll: Ensure IRQs stay disabled after cpuidle_state::enter() calls
...
Make cpuidle_state::enter() methods IRQ state invariant on exit.
Additionally make sure to use raw_local_irq_*() methods since this
cpuidle callback will be called with RCU already disabled.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Rafael J. Wysocki <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:21 +01:00
Peter Zijlstra
aaa3896b96
x86/idle: Replace 'x86_idle' function pointer with a static_call
...
Typical boot time setup; no need to suffer an indirect call for that.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Reviewed-by: Frederic Weisbecker <[email protected] >
Reviewed-by: Rafael J. Wysocki <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:21 +01:00
Peter Zijlstra
1f7c232ee0
x86/perf/amd: Remove tracing from perf_lopwr_cb()
...
The perf_lopwr_cb() function is called from the idle routines; there
is no RCU there, we must not enter tracing.
Use __always_inline, noidle annotations and existing no-trace methods.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Tested-by: Tony Lindgren <[email protected] >
Tested-by: Ulf Hansson <[email protected] >
Acked-by: Rafael J. Wysocki <[email protected] >
Acked-by: Frederic Weisbecker <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-13 11:03:21 +01:00
Peter Zijlstra
7c6dd961d0
x86/boot: Avoid using Intel mnemonics in AT&T syntax asm
...
With 'GNU assembler (GNU Binutils for Debian) 2.39.90.20221231' the
build now reports:
arch/x86/realmode/rm/../../boot/bioscall.S: Assembler messages:
arch/x86/realmode/rm/../../boot/bioscall.S:35: Warning: found `movsd'; assuming `movsl' was meant
arch/x86/realmode/rm/../../boot/bioscall.S:70: Warning: found `movsd'; assuming `movsl' was meant
arch/x86/boot/bioscall.S: Assembler messages:
arch/x86/boot/bioscall.S:35: Warning: found `movsd'; assuming `movsl' was meant
arch/x86/boot/bioscall.S:70: Warning: found `movsd'; assuming `movsl' was meant
Which is due to:
PR gas/29525
Note that with the dropped CMPSD and MOVSD Intel Syntax string insn
templates taking operands, mixed IsString/non-IsString template groups
(with memory operands) cannot occur anymore. With that
maybe_adjust_templates() becomes unnecessary (and is hence being
removed).
More details: https://sourceware.org/bugzilla/show_bug.cgi?id=29525
Borislav Petkov further explains:
" the particular problem here is is that the 'd' suffix is
"conflicting" in the sense that you can have SSE mnemonics like movsD %xmm...
and the same thing also for string ops (which is the case here) so apparently
the agreement in binutils land is to use the always accepted suffixes 'l' or 'q'
and phase out 'd' slowly... "
Fixes: 7a734e7dd9 ("x86, setup: "glove box" BIOS calls -- infrastructure")
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Ingo Molnar <[email protected] >
Acked-by: Borislav Petkov (AMD) <[email protected] >
Link: https://lore.kernel.org/r/[email protected]
2023-01-10 13:03:23 +01:00
Peter Zijlstra
526970be53
sh/mm: Fix pmd_t for real
...
Because typing is hard...
Fixes: 0862ff059c ("sh/mm: Make pmd_t similar to pte_t")
Reported-by: Guenter Roeck <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Signed-off-by: Linus Torvalds <[email protected] >
2023-01-10 05:31:42 -06:00
Peter Zijlstra
a551844e34
perf: Fix use-after-free in error path
...
The syscall error path has a use-after-free; put_pmu_ctx() will
reference ctx, therefore we must ensure ctx is destroyed after pmu_ctx
is.
Fixes: bd27568117 ("perf: Rewrite core context handling")
Reported-by: [email protected]
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Tested-by: Chengming Zhou <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-27 12:44:01 +01:00
Peter Zijlstra
e996365ee7
x86/mm: Rename __change_page_attr_set_clr(.checkalias)
...
Now that the checkalias functionality is taken by CPA_NO_CHECK_ALIAS
rename the argument to better match is remaining purpose: primary,
matching __change_page_attr().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221110125544.661001508%40infradead.org
2022-12-15 10:37:28 -08:00
Peter Zijlstra
d597416683
x86/mm: Inhibit _PAGE_NX changes from cpa_process_alias()
...
There is a cludge in change_page_attr_set_clr() that inhibits
propagating NX changes to the aliases (directmap and highmap) -- this
is a cludge twofold:
- it also inhibits the primary checks in __change_page_attr();
- it hard depends on single bit changes.
The introduction of set_memory_rox() triggered this last issue for
clearing both _PAGE_RW and _PAGE_NX.
Explicitly ignore _PAGE_NX in cpa_process_alias() instead.
Fixes: b38994948567 ("x86/mm: Implement native set_memory_rox()")
Reported-by: kernel test robot <[email protected] >
Debugged-by: Dave Hansen <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221110125544.594991716%40infradead.org
2022-12-15 10:37:28 -08:00
Peter Zijlstra
ef9ab81af6
x86/mm: Untangle __change_page_attr_set_clr(.checkalias)
...
The .checkalias argument to __change_page_attr_set_clr() is overloaded
and serves two different purposes:
- it inhibits the call to cpa_process_alias() -- as suggested by the
name; however,
- it also serves as 'primary' indicator for __change_page_attr()
( which in turn also serves as a recursion terminator for
cpa_process_alias() ).
Untangle these by extending the use of CPA_NO_CHECK_ALIAS to all
callsites that currently use .checkalias=0 for this purpose.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221110125544.527267183%40infradead.org
2022-12-15 10:37:28 -08:00
Peter Zijlstra
5ceeee7571
x86/mm: Add a few comments
...
It's a shame to hide useful comments in Changelogs, add some to the
code.
Shamelessly stolen from commit:
c40a56a781 ("x86/mm/init: Remove freed kernel image areas from alias mapping")
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221110125544.460677011%40infradead.org
2022-12-15 10:37:28 -08:00
Peter Zijlstra
2dff2c359e
mm: Convert __HAVE_ARCH_P..P_GET to the new style
...
Since __HAVE_ARCH_* style guards have been depricated in favour of
defining the function name onto itself, convert pxxp_get().
Suggested-by: Linus Torvalds <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:27 -08:00
Peter Zijlstra
eb780dcae0
mm: Remove pointless barrier() after pmdp_get_lockless()
...
pmdp_get_lockless() should itself imply any ordering required.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114425.298833095%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
d4a72e7fe6
x86/mm/pae: Get rid of set_64bit()
...
Recognise that set_64bit() is a special case of our previously
introduced pxx_xchg64(), so use that and get rid of set_64bit().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114425.233481884%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
9ee850acd2
x86_64: Remove pointless set_64bit() usage
...
The use of set_64bit() in X86_64 only code is pretty pointless, seeing
how it's a direct assignment. Remove all this nonsense.
[nathanchance: unbreak irte]
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114425.168036718%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
b7301f2010
x86/mm/pae: Be consistent with pXXp_get_and_clear()
...
Given that ptep_get_and_clear() uses cmpxchg8b, and that should be by
far the most common case, there's no point in having an optimized
variant for pmd/pud.
Introduce the pxx_xchg64() helper to implement the common logic once.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114425.103392961%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
f7bcd4617d
x86/mm/pae: Use WRITE_ONCE()
...
Disallow write-tearing, that would be really unfortunate.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114425.038102604%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
7a9b8bdb6a
x86/mm/pae: Don't (ab)use atomic64
...
PAE implies CX8, write readable code.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.971450128%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
1180e732c9
mm/gup: Fix the lockless PMD access
...
On architectures where the PTE/PMD is larger than the native word size
(i386-PAE for example), READ_ONCE() can do the wrong thing. Use
pmdp_get_lockless() just like we use ptep_get_lockless().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.906110403%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
dab6e71742
mm: Rename pmd_read_atomic()
...
There's no point in having the identical routines for PTE/PMD have
different names.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.841277397%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
6ca297d478
mm: Rename GUP_GET_PTE_LOW_HIGH
...
Since it no longer applies to only PTEs, rename it to PXX.
Suggested-by: Linus Torvalds <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.776404066%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
024d232ae4
mm: Fix pmd_read_atomic()
...
AFAICT there's no reason to do anything different than what we do for
PTEs. Make it so (also affects SH).
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.711181252%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
0862ff059c
sh/mm: Make pmd_t similar to pte_t
...
Just like 64bit pte_t, have a low/high split in pmd_t.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.645657294%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
fbfdec9989
x86/mm/pae: Make pmd_t similar to pte_t
...
Instead of mucking about with at least 2 different ways of fudging
it, do the same thing we do for pte_t.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.580310787%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
93b3037a14
mm: Update ptep_get_lockless()'s comment
...
Improve the comment.
Suggested-by: Matthew Wilcox <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/20221022114424.515572025%40infradead.org
2022-12-15 10:37:27 -08:00
Peter Zijlstra
60463628c9
x86/mm: Implement native set_memory_rox()
...
Provide a native implementation of set_memory_rox(), avoiding the
double set_memory_ro();set_memory_x(); calls.
Suggested-by: Linus Torvalds <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
2022-12-15 10:37:27 -08:00
Peter Zijlstra
d48567c9a0
mm: Introduce set_memory_rox()
...
Because endlessly repeating:
set_memory_ro()
set_memory_x()
is getting tedious.
Suggested-by: Linus Torvalds <[email protected] >
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:26 -08:00
Peter Zijlstra
414ebf148c
x86/mm: Do verify W^X at boot up
...
Straight up revert of commit:
a970174d7a ("x86/mm: Do not verify W^X at boot up")
now that the root cause has been fixed.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:26 -08:00
Peter Zijlstra
eb7d389d5b
x86/ftrace: Remove SYSTEM_BOOTING exceptions
...
Now that text_poke is available before ftrace, remove the
SYSTEM_BOOTING exceptions.
Specifically, this cures a W+X case during boot.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:26 -08:00
Peter Zijlstra
5b93a83649
x86/mm: Initialize text poking earlier
...
Move poking_init() up a bunch; specifically move it right after
mm_init() which is right before ftrace_init().
This will allow simplifying ftrace text poking which currently has
a bunch of exceptions for early boot.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:26 -08:00
Peter Zijlstra
3f4c8211d9
x86/mm: Use mm_alloc() in poking_init()
...
Instead of duplicating init_mm, allocate a fresh mm. The advantage is
that mm_alloc() has much simpler dependencies. Additionally it makes
more conceptual sense, init_mm has no (and must not have) user state
to duplicate.
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:26 -08:00
Peter Zijlstra
af80602799
mm: Move mm_cachep initialization to mm_init()
...
In order to allow using mm_alloc() much earlier, move initializing
mm_cachep into mm_init().
Signed-off-by: Peter Zijlstra (Intel) <[email protected] >
Link: https://lkml.kernel.org/r/[email protected]
2022-12-15 10:37:26 -08:00