The CBC ciphers require a fully unpredictable initialisation vector,
which we currently generate as a channel ephemeral secret. The GCM
ciphers require only a unique initialisation vector: there is no
requirement for it also to be unpredictable. The content of the
record IV portion of the IV is a free choice of the sender, and we
currently use an unpredictable value for both CBC and GCM.
TLS version 1.3 removes the record IV portion for GCM ciphers, instead
constructing the IV by XORing the sequence number into the end of the
fixed IV.
Define the concept of a sequential initialisation vector as meaning
that the sequence number is XORed into the end of the overall
initialisation vector (which may be either the fixed IV or the record
IV portion), with no per-record unpredictable value required. This
allows us to represent the mechanism required for TLS version 1.3, and
avoid the unnecessary cost of generating a channel ephemeral secret
for a GCM cipher under TLS version 1.2.
On the receive side, the XORed portion may be overwritten by the real
record IV, since the sender's choice is always definitive for the
contents of the record IV.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
TLS version 1.3 masquerades as TLS version 1.2 on the wire for the
benefit of badly engineered firewalls and other intermediate devices
that attempt to inspect the protocol stream.
Limit the maximum version in transmitted record headers to be TLS
version 1.2. (Continue to accept any version in received record
headers, since this value has never had any significance.)
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The cipher specifications are currently modelled as an active cipher
specification that corresponds to the cipher currently in use by the
secure channel abstraction, and a pending cipher specification that
corresponds to the cipher that will be swapped in after the next
ChangeCipherSpec.
This design reflects the wording of RFC 2246 through to RFC 5246:
"there are always four connection states outstanding: the current read
and write states, and the pending read and write states".
This model does not map well to TLS version 1.3, with its multiple
phases of traffic secrets and somewhat idiosyncratic choices of
transcript boundaries. The client Finished message is a particular
problem: the client application traffic secret must be calculated
after constructing the client Finished verify_data but before adding
the client Finished to the transcript digest (i.e. before encrypting
it with the client handshake traffic keys). This is an irritating
asymmetry with the server application traffic secret, which may be
calculated cleanly after the server Finished message has been added to
the transcript digest.
Switch to a model in which only the active cipher specification
exists, and always corresponds to the cipher currently in use by the
secure channel abstraction.
Move the record sequence number to become part of the cipher
specification, so that the sequence number reset logic can be shared
between the transmit and receive paths.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
RFC 8446 defines the key_share extension as a mechanism for ephemeral
key exchange (replacing ServerKeyExchange and ClientKeyExchange).
Add support for sharing a key in our ClientHello and for parsing the
shared key from a ServerHello. Construct the ClientHello on the heap
rather than on the stack, since the shared key values may be large
(e.g. for FFDHE4096).
Select the most preferred named group for sending the initial shared
key. (We do not yet handle a HelloRetryRequest: if the server chooses
a different named group then we will record this group but do not yet
support sending the second ClientHello.)
We do not explicitly reject a key share extension received from a
server that negotiated TLS version 1.2 or lower. Any such key will be
successfully used to establish a shared secret (and so the secure
channel will become keyed), but there is no way for this shared secret
to subsequently be successfully bound to the server's identity: a
ServerKeyExchange would replace the shared secret (since the TLS
version 1.2 key schedule is not accumulative), and a CertificateVerify
would fail to generate a signable digest.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The TLS version 1.3 ClientHello includes a key_share extension whose
value will depend upon the selected key exchange named group. The
incorporation of the initial ClientHello into the selected handshake
digest must therefore be done before the named group is potentially
modified.
Move responsibility for adding the initial ClientHello to the
transcript digest from tls_new_server_hello() to tls_select_cipher(),
so that this can be done before updating the selected named group.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
RFC 8446 redefines the session ID within a TLS version 1.3 ServerHello
as being a field that must always echo the session ID from ClientHello
(and no longer indicates that session resumption is taking place), as
a workaround for badly engineered middleware boxes.
Skip ID-based session resumption if the negotiated version is TLS
version 1.3 or later, and instead abort the connection if the session
ID is not echoed verbatim (as per the RFC).
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The signable digest value is used to bind the server identity to the
shared secret, and so the digest must be computed over the parameters
used to establish the shared secret.
For both endpoints in TLS version 1.3 and for the client endpoint in
TLS version 1.2, the digest is computed over the running transcript
hash and so already includes the parameters used to establish the
shared secret.
For the server endpoint in TLS version 1.2, the digest is computed
over only the client and server random bytes plus any additional data
passed in by the caller, and so this additional data must include the
parameters used to establish the shared secret.
The signable digest value for the server endpoint is currently
computed only in response to a ServerKeyExchange record, in which case
the additional data correctly contains the parameters from that record
that were used to establish the shared secret.
Adding support for TLS version 1.3 will necessitate adding the ability
to parse a received CertificateVerify record, which will attempt to
construct a signable digest value with no additional data.
Require additional data to be provided when constructing a signable
digest value for the server endpoint using the TLS version 1.2 key
schedule, to prevent a CertificateVerify from potentially being used
to verify a digest that was not computed over the parameters used to
establish the shared secret.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The signable digest value is used to bind the server identity to the
shared secret, and so the digest must be computed over the parameters
used to establish the shared secret in order to be meaningful.
The secure channel will refuse to bind the peer identity on the basis
of a verified signable digest if the channel does not already contain
key material derived from a shared secret.
A signable digest that was erroneously constructed before a shared
secret was applied is therefore guaranteed to be unusable for binding
the channel, provided that the caller uses a sensible sequence of
operations (i.e. constructs the signable digest and then immediately
attempts to use it to bind the peer identity).
Strengthen this guarantee further by refusing to generate a signable
digest value unless the key schedule already contains key material.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
RFC 8446 defines a mechanism that allows (but does not guarantee) the
detection of version downgrade attacks, based on magic signature
values placed within the ServerHello random bytes. The magic
signature will be present if the server supports any version higher
than the negotiated version, and so may be present if the server
supports a higher version than we are offering.
The last byte of the magic signature is non-constant and is defined to
match the server's negotiated protocol version, encoded as a delta
from the value 0x0302 representing TLS version 1.1 (or lower). Since
the server random bytes are always used in the construction of
verify_data (even in older versions of TLS without the extended master
secret), this encoding of the negotiated version cannot be forged by
an attacker.
If the version that is negotiated is lower than the version that we
offered (i.e. if a downgrade attack could possibly be happening), then
check for the range of magic signatures that could indicate a
downgrade attack, and terminate the connection if applicable.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
RFC 8446 caps the version number field in ClientHello and ServerHello
to represent at most TLS version 1.2, as a workaround for badly
engineered servers and middleware boxes. The actual protocol version
is instead negotiated via the supported_versions extension.
Send the list of supported versions in the ClientHello, and parse the
selected version from the supported_version extenion if present in the
ServerHello.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
RFC 8446 redefines the concept of a cipher suite for TLS version 1.3
to exclude the key exchange algorithm, leaving it specifying only the
block cipher algorithm and the handshake digest algorithm.
Add definitions for the two cipher suites that we can currently
support (TLS_AES_128_GCM_SHA256 and TLS_AES_256_GCM_SHA384).
We define these as using the null key exchange algorithm. The null
key exchange algorithm will fail on any attempt at key agreement. A
server that attempts to rely on the key exchange algorithm implied by
the cipher suite (e.g. a server attempting to illegally use these
cipher suites with TLS version 1.2) will therefore be unable to
establish a shared secret and so will not be able to cause the secure
channel to become established.
We therefore do not explicitly check for and reject a server's attempt
to negotiate a TLS version 1.3 cipher suite under TLS version 1.2 or
earlier: the secure channel abstraction already ensures that such a
negotiation is doomed to failure.
Under TLS version 1.3, the cipher suite's key exchange algorithm
specification will not be used. We therefore do not explicitly check
for and reject a server's attempt to negotiate a TLS version 1.2 or
earlier cipher suite under TLS version 1.3 or later: we instead just
ignore the key exchange algorithm aspect of that cipher suite.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The concept of a named group is currently relevant only at the point
of parsing a ServerKeyExchange record to determine the key exchange
algorithm: once parsed, the group's numeric code is no longer required
and so we currently record only the resulting key exchange algorithm.
For TLS version 1.3, the numeric code will also be needed when
constructing the key_share extension in the ClientHello.
Switch from recording the key exchange algorithm to recording the
functionally equivalent named group, thereby making it possible to
retrieve the numeric code when needed.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Implement AES hardware acceleration for AArch64 using the AES subset
of the Cryptographic Extensions. All supported runtime environments
already allow for use of the SIMD/FP registers, and so the only
required compiler quirk is to annotate the functions as being
permitted to emit the AES instructions.
Feature detection relies upon the ability to read the ID_AA64ISAR0_EL1
system register. This works as expected in all supported runtime
environments:
- As a UEFI binary, we are running in a real EL1 and so can just
read the system register for the current (and only active) core
- As a Linux userspace binary running in EL0, the kernel (since
4.11) will emulate the read to report the subset of features that
are supported by all online cores
- As a Linux userspace binary run via QEMU's binary translation,
QEMU (since 4.0.0) will similarly emulate the read to report the
features supported by the selected CPU model
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Implement AES hardware acceleration for i386 and x86_64, using the
AES-NI instructions (and the SSE2 "pxor" instruction for the initial
AddRoundKey), using the unmodified existing key schedule as generated
by aes_setkey().
For raw AES (ignoring the block cipher mode of operation), this
results in a speed improvement from approximately 20 cycles per byte
down to approximately 1 cycle per byte.
The AES-NI instructions use SSE registers. For the sake of not having
to think about the possible consequences across all various runtime
environments (BIOS/UEFI/Linux), we choose not to enable "-msse" in
CFLAGS for this file. We include ".arch" directives to ensure that
the assembler knows that it is permitted to emit the SSE2 and AES-NI
instructions, use a fixed "%xmm0" rather than an "x" constraint (which
GCC would consider to be impossible without "-msse"), and restore the
value of "%xmm0" after use to meet the requirements of the most
restrictive ABI for which this file can be built.
In an ideal world, we would also use a ".arch push" / ".arch pop" pair
to restore the permitted instruction set, rather than leaving the SSE2
and AES-NI instructions as permitted outside the scope of the inline
asm. Unfortunately this feature would require binutils 2.34 or newer,
and so would prevent building iPXE on some still-current distros such
as RHEL8.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Some AES hardware acceleration instructions require each 16-byte round
key to be aligned on a 16-byte boundary. Cipher contexts are byte
arrays allocated by the caller and do not have any guaranteed
alignment.
Increase the AES context size to allow space for alignment padding,
and align the context before use. Reduce the round count field from
an unsigned int to a uint8_t, to minimise wasted space and to ensure
that the resulting padded context size is itself reasonably aligned.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Allow architectures to detect support for AES hardware acceleration at
runtime and to replace the AES algorithm's encrypt() and decrypt()
method pointers with hardware accelerated implementations.
Extend the automated tests to run the AES tests twice: once with
hardware acceleration explicitly disabled (to test the unaccelerated
software implementation) and once with acceleration re-enabled. Skip
the second test if no hardware acceleration is available: this avoids
unnecessarily repeating the test of the unaccelerated implementation,
and allows a non-zero test count for "aes-hw" to indicate that the
hardware acceleration was tested. For example:
On a system that supports AES hardware acceleration:
OK: "aes" 120 tests passed
OK: "aes-hw" 120 tests passed
On a system that does not support AES hardware acceleration:
OK: "aes" 120 tests passed
OK: "aes-hw" 0 tests passed
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Clear CR0.EM, clear CR0.TS, and set CR4.OSFXSR in order to allow the
execution of SSE instructions such as the AES-NI instructions for AES
hardware acceleration.
Note that we have to assert CR4.OSFXSR to allow these instructions to
execute, but our context switching logic (e.g. the protected-mode and
long-mode interrupt handlers) does not actually preserve the
FPU/MMX/SSE registers.
Our C code therefore cannot in general presume that the FPU/MMX/SSE
registers will be preserved across arbitrary context boundaries.
However, since C code executes with interrupts disabled, an individual
function may safely use temporary FPU/MMX/SSE registers provided that
it does not enable interrupts or otherwise relinquish the context.
For the same C code to also be usable under the UEFI IA-32 ABI, it
must preserve all registers other than %eax, %ecx, and %edx, including
preserving all MMX and XMM registers.
The practical upshot is therefore that C code may use SSE instructions
and may assume that SSE registers will not be changed arbitrarily
during execution (either because the ABI guarantees preservation, as
with UEFI or Linux, or because the runtime environment guarantees that
interrupts are disabled), but the C code must itself restore the
values of any modified FPU/MMX/SSE registers.
The "fxsave"/"fxrstor" performed by virt_call() would allow for a
slightly more relaxed constraint if support for the UEFI IA-32 ABI
were ever to be dropped in future. A real-mode caller that is making
use of SSE must have already set OSFXSR, and so its non-64-bit
registers %xmm0-%xmm7 would already be saved and restored across
virt_call(). The tightest constraint would then become the UEFI X64
ABI, which defines %xmm0-%xmm5 as volatile (i.e. caller-saved): this
would allow C code in iPXE to use %xmm0-%xmm5 without needing to
explicitly save and restore their values.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Clearing the CR0.EM and CR0.TS flags is a prerequisite for using the
AES-NI instructions for AES hardware acceleration: if CR0.EM is set
then the CPU will raise an undefined-instruction exception, and if
CR0.TS is set then the CPU will raise a device-not-available exception
(expecting the OS to have installed an exception handler that would
perform a deferred context switch of the FPU/MMX/SSE registers).
Preserve CR0 across virt_call(), to allow the CR0.EM and CR0.TS flags
to be modified as needed.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Setting the CR4.OSFXSR flag is a prerequisite for using the AES-NI
instructions for AES hardware acceleration: if this flag is not set
then the CPU will raise an undefined-instruction exception.
We currently preserve CR4 across virt_call() only for 64-bit builds,
since those will modify CR4 by setting CR4.PAE. Extend this to
preserve CR4 across virt_call() if FXSR is supported, to allow the
CR4.OSFXSR flag to be modified as needed.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
SSE is an architectural requirement for x86_64, and FXSR is an
architectural requirement for SSE. We can therefore skip the FXSR
check in a 64-bit build, since no 64-bit CPU can exist that does not
advertise support for FXSR.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Commit 71560d1 ("[librm] Preserve FPU, MMX and SSE state across calls
to virt_call()") originally introduced the use of "fxsave" and
"fxrstor" to work around a bug in the implementation of memcpy()
within the IBM Tivoli Provisioning Manager's VMM. This commit assumed
(with justification given in the commit message) that SSE support
could be assumed to be present on any realistic in-scope CPU, and so
these instructions may safely be assumed to be supported.
Commit dd9a14d ("[librm] Conditionalize the workaround for the Tivoli
VMM's SSE garbling") then made this workaround a compile-time
conditional, to work around a missing feature in QEMU that caused the
use of "fxsave" and "fxrstor" to fail in QEMU VMs on some host CPUs.
Commit 900f1f9 ("[librm] Test for FXSAVE/FXRSTOR instruction support")
then added a runtime CPUID check for the FXSR feature, to allow the
unmodified iPXE binary to be used on older CPUs.
Supporting AES hardware acceleration via AES-NI will require setting
the CR4.OSFXSR control bit, which in turn must be conditionalised upon
the same runtime CPUID check for the FXSR feature.
Perform the runtime check unconditionally, and assume that we no
longer need the compile-time conditional (i.e. assume either that
newer versions of QEMU emulate "fxsave" and "fxrstor" when needed, or
that QEMU reports via CPUID that FXSR is not supported if it cannot
support those instructions).
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The separate VC_TMP_GDT and VC_TMP_IDT fields suggest that these could
be moved freely relative to each other. This is not the case: callers
of prot_to_real() must pass a single pointer to the combined pair.
Collapse to a single field, with the name adjusted to VC_TMP_GDTR_IDTR
to more closely match the rm_default_gdtr_idtr structure that
necessarily shares the same layout.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Commit 6143057 ("[librm] Add support for running in 64-bit long mode")
treated disabling paging on the transition into protected mode as
something that needed to be done as a precaution only in a 64-bit
build, on the assumption that in a 32-bit BIOS system nothing else
would be enabling paging.
The Intel SDM states that setting CR0.PG in real mode (with CR0.PE
clear) will raise a general-protection exception anyway, and so we
should never encounter a situation in which CR0.PG is set at this
point. A review of the bochs source code suggests that it may be
possible to encounter the combination of CR0.PG set with CR0.PE clear
in an SVM guest.
Err on the side of paranoia and always disable paging as part of the
transition from real mode to protected mode.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
ASN.1 allows for multi-byte tag numbers by setting the low five bits
of the first tag byte to 0x1f. No tag that we need to handle has this
format, and the existing checks for specific tag numbers will already
fail to match against such a tag (treating it as a normal single-byte
tag number).
Refuse to parse any tag with a high tag number format, to guard
against future bugs that could arise because the tag length would be
calculated incorrectly.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The Project Wycheproof self-tests are deliberately not included in the
normal per-commit test suite since they are extremely slow to run.
Add a build target that includes the slow self-tests, and run these
tests on a push to the "slowtest" branch.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
As detailed in commit 511dfd2 ("[crypto] Reject non-canonical ECDSA
signature data structures"), changing the representation of a valid
ECDSA signature to a different valid representation of the same
signature does not conceptually make it an invalid signature.
However, some large public test vector sets conflate the concepts of
"altered representation" and "invalid representation" in a way that
makes it difficult to determine which tests ought to pass and which
ought to fail without extensive manual analysis.
Reject any ECDSA signature object that does not have the expected
total length. The signature parsing logic already ensures that the
expected structure exists, and so the total length can be correct only
if every object used the expected DER encoding.
This length check completely subsumes the checks that were introduced
in commit 511dfd2 ("[crypto] Reject non-canonical ECDSA signature data
structures"), since there is no way to insert additional information
without also affecting the length.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
An indefinite length encoding will currently be parsed as having a
length of zero, and an encoded length that exceeds the range of an
unsigned int will be truncated.
Tighten up the parsing of lengths to explicitly reject indefinite
length encodings or unrepresentable lengths.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
We use asn1_enter_unsigned() essentially as a convenience mechanism to
skip the initial zero byte found when an encoder had to insert the
zero to prevent a logically unsigned value from being interpreted as
negative.
We currently accept malformed values where the initial byte has the
MSB set, and allow them to be interpreted as unsigned values.
Tighten up the parsing of unsigned integers so that values where the
initial byte has the MSB set will be rejected as invalid, and ensure
that the resulting cursor is minimal by skipping any number of initial
padding zero bytes (so that asn1_compare() can then be used without
the risk of false negatives).
Signed-off-by: Michael Brown <mcb30@ipxe.org>
An ECDSA signature value is a vector of two integers (r,s) modulo the
curve group order. The ECDSA algorithm itself does not define the
encoding to be used for these two integers. At least two different
standards exist for representing the vector (r,s): the ASN.1 structure
originally defined in RFC 3279 (which uses a SEQUENCE of two INTEGER
values) and the raw byte concatenation structure defined in IEEE
P1363. A valid signature vector (r,s) may be freely converted between
these two formats. Changing the format does not logically change the
validity of the signature.
Due to the mathematics underlying ECDSA, the vector (r,-s) is also
always a valid signature for the same content.
With the ASN.1 structure, there exists the possibility of adding extra
data that would currently be ignored by the parser: either objects
following the top-level SEQUENCE, or objects within the SEQUENCE
following the two INTEGER values. Adding this data does not logically
change the validity of the signature, in the same way that converting
between ASN.1 and P1363 does not logically change the validity of the
signature. However, some public test vector sets check for the
rejection of signatures containing inserted data.
Reject any ECDSA signature object that includes data following the
top-level SEQUENCE, or that includes data following the "r" and "s"
INTEGER values.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
We don't need the ability to validate that the test flags are
appropriate for the type of test. Reduce duplication by using a
single shared class for all existent test flags.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
An RSA signature value (or encrypted message value) is a congruence
class modulo the field prime, and so adding a multiple of the field
prime does not logically change the validity of the signature (or the
content of the encrypted message).
However, RFC 8017 states that both the decryption primitive (section
5.1.2) and the verification primitive (section 5.2.2) should reject
non-canonical input values (i.e. any value that is not strictly less
than the field prime), and some public test vector sets check for this
rejection.
Treat any signature value or encrypted message value that is equal to
or greater than the field prime as being invalid.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
RSA PKCS#1 requires a minimum of eight non-zero padding bytes for
encryption. This limit is currently enforced when encrypting but not
validated when decrypting.
The only existing code path that can currently lead to RSA decryption
is CMS decryption, which can use RSA to decrypt the cipher key. With
underlength PKCS#1 padding (and hence an overlength plaintext), the
decrypted cipher key would be rejected by the immediately following
call to cipher_setkey().
Add the missing padding length check as part of RSA decryption, and
add a test case (imported from Project Wycheproof) to ensure that this
check remains in place in future.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
Some public RSA and ECDSA test vector sets provide only the private or
public half of the key pair, and reuse the same key for multiple tests
within the set.
Allow public-key tests to omit either half of the key pair, and to
therefore perform separate tests for encryption, decryption, signature
generation, and signature verification.
The existing RSA and ECDSA tests (which all include a full key pair)
are unchanged by this reorganisation.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The Project Wycheproof RSA and ECDSA tests define the key as a
property of the test group rather than of the individual test case.
Allow test groups to have stable identifiers and to participate in
generating the source code for the test definitions.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
The schema and test counts are currently validated only at the point
of attempting to generate source code. Promote these checks to become
standard validators.
Signed-off-by: Michael Brown <mcb30@ipxe.org>
A zero-length IV is not permitted by the NIST GCM specification, since
it would lead to leaking the authentication key.
The only existing code path that can currently lead to the use of a
GCM cipher with a zero-length initialisation vector is CMS decryption.
Modifying a CMS encrypted message to include a zero-length IV would
leak information required to obtain the authentication key into the
transient decrypted image, but this transient image would then fail
the GCM authentication tag check and so the decrypted plaintext would
be immediately overwritten (with the re-encrypted ciphertext).
Improve robustness by rejecting a zero-length initialisation vector
for a GCM cipher, and add a test case to ensure that this rejection
remains in place in future.
Signed-off-by: Michael Brown <mcb30@ipxe.org>