From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail-oo1-xc33.google.com (mail-oo1-xc33.google.com [IPv6:2607:f8b0:4864:20::c33]) by sourceware.org (Postfix) with ESMTPS id 1F9673858C2C for ; Wed, 20 Dec 2023 02:04:15 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 1F9673858C2C Authentication-Results: sourceware.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=gmail.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 1F9673858C2C Authentication-Results: server2.sourceware.org; arc=none smtp.remote-ip=2607:f8b0:4864:20::c33 ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1703037857; cv=none; b=uENk9kVgaCP0ZAhdTd6OFSd0TfffZOWW1b3h3saG5cjRJbKhJHHjGzimucp5ubb9Bwo1Qy7NzbAbQlqxUjvcIhAgpCx5E93Tf+Z8NsCoQbqH3SOk+2hfAiTPrQ2Zj2uRbja8VFCbDkpiXTx/YJvlXbUYomT70htRClM7uKtI8wM= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1703037857; c=relaxed/simple; bh=s5vn3NFaS7gxz5SwCqvKLCyyJ199c+NErPHrGJUci6M=; h=DKIM-Signature:MIME-Version:From:Date:Message-ID:Subject:To; b=THWVXWLWl/DdWDiwfUlcapAd/+JrsXlEU31pwl28+hTBaD6pN0b3FOUX/DR9UZkjNkpnvkirx4lpv+pj61omTQVYVWLdtK1EVgCs9chtAK/Cas81g3N+NpWMB5SfSFxNV4gOLoQUM3wGAcDts4/13LTTPMv2fx7LpHYI/2XAsNA= ARC-Authentication-Results: i=1; server2.sourceware.org Received: by mail-oo1-xc33.google.com with SMTP id 006d021491bc7-593fa46fd45so504977eaf.2 for ; Tue, 19 Dec 2023 18:04:15 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1703037854; x=1703642654; darn=gcc.gnu.org; h=content-transfer-encoding:to:subject:message-id:date:from :in-reply-to:references:mime-version:from:to:cc:subject:date :message-id:reply-to; bh=HrtFfxnHqD/xdC+pAm6wRAhJhRopFTtgXamaAIY6++4=; b=mhjBuzq4CLk9Q/yVnqmuupmh6HMM3s/vVhDV62Sf8gY48g/7shobg6mjHxCcFX2TAZ bY82nGiPSq5uLwThkDWLWp3UNU2G2gAfkdHsobSyINn70BHyNkKQjjrV9+H143fJtjql wWMpZNLKm4Egi5Gys6/VmFC1tGkFawxKVHdzTAFSBRNdeSA2qB0wonkql0zsPrsJFe1L OWUush6FqspdlceeXHXDrZvdGW7t+1EGajy7/Y805gx1+0lJJJJKFC2WokFLpV85VPGt S1RNaAa3KkpQqhKXm9cXW2uMB0sX58JYwnUcZCkEH6IYdS0n/dTwQBn7ytz4t0bl3zYR rbuA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1703037854; x=1703642654; h=content-transfer-encoding:to:subject:message-id:date:from :in-reply-to:references:mime-version:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=HrtFfxnHqD/xdC+pAm6wRAhJhRopFTtgXamaAIY6++4=; b=VRoKyWQ1UcBS1tjui44o+yju58Du7yGPGUSsTkHqAfmdIls3O5eFMa40ZLpns/Br4w e1vgNptnizD7ZVUKMQor/Ka3i8lI8zzA+T9e2CEtEYsiTO9Rkcevk1RK1UPxXSr6RELR tDne5v83DQoh+N3TEUmO1rZk154pwqow2Y5VMNMlwblFYT8ci7OF48Y8xI2QoPd2qkVA W5TdKnGbcCUSFUGM+eX+Y/3tW8V7UJhChhZWQooqSrP852aF+VkaHIZqNb6Siw9ruNj9 r24pUJ++y498YW5a1IHl6mvjbMapoZE3GOekFE9NGk04nNWtzXCpGgts/aTUabjK989J A/fQ== X-Gm-Message-State: AOJu0YyquKiwA71oJyv/hvJIplPXDe4//Ul589gAyiWKXcER1dEzcKwr NODT92KPKPqWn7dcJAEZApVoMJJ832OEIpw2y0k= X-Google-Smtp-Source: AGHT+IGnOYZLOiPdzw6JJfFer+bFANbGnW6UcY+j/iXAxAJUmoYce2hODBfH+yiDpBj0IY/AbZxfAVWZhKGaGajzeP8= X-Received: by 2002:a05:6358:a088:b0:172:ae52:ac40 with SMTP id u8-20020a056358a08800b00172ae52ac40mr6059606rwn.38.1703037854146; Tue, 19 Dec 2023 18:04:14 -0800 (PST) MIME-Version: 1.0 References: <097AABD6596FB0C3+2023121906491281154423@rivai.ai> <92p02r8p-46rq-s976-5r8p-s87q0q763465@fhfr.qr> <6FD0A43E2F3E9BD9+202312191735136921653@rivai.ai> In-Reply-To: From: Andrew Pinski Date: Tue, 19 Dec 2023 18:04:01 -0800 Message-ID: Subject: Re: [PATCH] fold-const: Handle AND, IOR, XOR with stepped vectors [PR112971]. To: Richard Biener , "juzhe.zhong@rivai.ai" , Robin Dapp , gcc-patches , "pan2.li" , Richard Biener , pinskia , richard.sandiford@arm.com Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Spam-Status: No, score=-8.4 required=5.0 tests=BAYES_00,DKIM_SIGNED,DKIM_VALID,DKIM_VALID_AU,DKIM_VALID_EF,FREEMAIL_FROM,GIT_PATCH_0,KAM_MANYTO,KAM_SHORT,RCVD_IN_DNSWL_NONE,SPF_HELO_NONE,SPF_PASS,TXREP,T_SCC_BODY_TEXT_LINE autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on server2.sourceware.org List-Id: On Tue, Dec 19, 2023 at 2:40=E2=80=AFAM Richard Sandiford wrote: > > Richard Biener writes: > > On Tue, 19 Dec 2023, juzhe.zhong@rivai.ai wrote: > > > >> Hi, Richard. > >> > >> After investigating the codes: > >> /* Return true if EXPR is the integer constant zero or a complex const= ant > >> of zero, or a location wrapper for such a constant. */ > >> > >> bool > >> integer_zerop (const_tree expr) > >> { > >> STRIP_ANY_LOCATION_WRAPPER (expr); > >> > >> switch (TREE_CODE (expr)) > >> { > >> case INTEGER_CST: > >> return wi::to_wide (expr) =3D=3D 0; > >> case COMPLEX_CST: > >> return (integer_zerop (TREE_REALPART (expr)) > >> && integer_zerop (TREE_IMAGPART (expr))); > >> case VECTOR_CST: > >> return (VECTOR_CST_NPATTERNS (expr) =3D=3D 1 > >> && VECTOR_CST_DUPLICATE_P (expr) > >> && integer_zerop (VECTOR_CST_ENCODED_ELT (expr, 0))); > >> default: > >> return false; > >> } > >> } > >> > >> I wonder whether we can simplify the codes as follows :? > >> if (integer_zerop (arg1) || integer_zerop (arg2)) > >> step_ok_p =3D (code =3D=3D BIT_AND_EXPR || code =3D=3D BIT_IOR= _EXPR > >> || code =3D=3D BIT_XOR_EXPR); > > > > Possibly. I'll let Richard S. comment on the whole structure. > > The current code is handling cases that require elementwise arithmetic. > ISTM that what we're really doing here is identifying cases where > whole-vector arithmetic is possible instead. I think that should be > a separate pre-step, rather than integrated into the current code. > > Largely this would consist of writing out match.pd-style folds in > C++ code, so Andrew's fix in comment 7 seems neater to me. I didn't like the change to match.pd (even with a comment on why) because it violates the whole idea behind canonicalization of constants being 2nd operand of commutative and comparison expressions. Maybe there are only a few limited match/simplify patterns which need to add the :c for constants not being the 2nd operand but there needs to be a comment on why :c is needed for this. > > But if this must happen in const_binop instead, then we could have > a function like: The reasoning of why it should be in const_binop rather than in match is because both operands are constants. Now for commutative expressions, we only need to check the first operand for zero/all_ones and try again swapping the operands. This will most likely solve the problem of writing so much code. We could even use lambdas to simplify things too. Thanks, Andrew Pinski > > /* OP is the INDEXth operand to CODE (counting from zero) and OTHER_OP > is the other operand. Try to use the value of OP to simplify the > operation in one step, without having to process individual elements. = */ > tree > simplify_const_binop (tree_code code, rtx op, rtx other_op, int index) > { > ... > } > > Thanks, > Richard > > > > > Richard. > > > >> > >> > >> > >> juzhe.zhong@rivai.ai > >> > >> From: Richard Biener > >> Date: 2023-12-19 17:12 > >> To: juzhe.zhong@rivai.ai > >> CC: Robin Dapp; gcc-patches; pan2.li; richard.sandiford; Richard Biene= r; pinskia > >> Subject: Re: Re: [PATCH] fold-const: Handle AND, IOR, XOR with stepped= vectors [PR112971]. > >> On Tue, 19 Dec 2023, juzhe.zhong@rivai.ai wrote: > >> > >> > Hi?Richard. Do you mean add the check as follows ? > >> > > >> > if (VECTOR_CST_NELTS_PER_PATTERN (arg1) =3D=3D 1 > >> > && VECTOR_CST_NELTS_PER_PATTERN (arg2) =3D=3D 3 > >> > >> Or <=3D 3 which would allow combining. As said, not sure what > >> =3D=3D 2 would be and whether that would work. > >> > >> Btw, integer_allonesp should also allow to be optimized for > >> and/ior at least. Possibly IOR/AND with the sign bit for > >> signed elements as well. > >> > >> I wonder if there's a programmatic way to identify OK cases > >> rather than enumerating them. > >> > >> > && integer_zerop (VECTOR_CST_ELT (arg1, 0))) > >> > step_ok_p =3D (code =3D=3D BIT_AND_EXPR || code =3D=3D BIT_I= OR_EXPR > >> > || code =3D=3D BIT_XOR_EXPR); > >> > else if (VECTOR_CST_NELTS_PER_PATTERN (arg2) =3D=3D 1 > >> > && VECTOR_CST_NELTS_PER_PATTERN (arg1) =3D=3D 3 > >> > && integer_zerop (VECTOR_CST_ELT (arg2, 0))) > >> > step_ok_p =3D (code =3D=3D BIT_AND_EXPR || code =3D=3D BIT_I= OR_EXPR > >> > || code =3D=3D BIT_XOR_EXPR); > >> > > >> > > >> > > >> > juzhe.zhong@rivai.ai > >> > > >> > From: Richard Biener > >> > Date: 2023-12-19 16:15 > >> > To: ??? > >> > CC: rdapp.gcc; gcc-patches; pan2.li; richard.sandiford; richard.guen= ther; Andrew Pinski > >> > Subject: Re: [PATCH] fold-const: Handle AND, IOR, XOR with stepped v= ectors [PR112971]. > >> > On Tue, 19 Dec 2023, ??? wrote: > >> > > >> > > Thanks Robin send initial patch to fix this ICE bug. > >> > > > >> > > CC to Richard S, Richard B, and Andrew. > >> > > >> > Just one comment, it seems that VECTOR_CST_STEPPED_P should > >> > implicitly include VECTOR_CST_DUPLICATE_P since it would be > >> > a step of zero (but as implemented it doesn't catch this). > >> > Looking at the implementation it's odd that we can handle > >> > VECTOR_CST_NELTS_PER_PATTERN =3D=3D 1 (duplicate) and > >> > =3D=3D 3 (stepped) but not =3D=3D 2 (not sure what that would be). > >> > > >> > Maybe the tests can be re-formulated in terms of > >> > VECTOR_CST_NELTS_PER_PATTERN? > >> > > >> > Richard. > >> > > >> > > Thanks. > >> > > > >> > > > >> > > > >> > > juzhe.zhong@rivai.ai > >> > > > >> > > From: Robin Dapp > >> > > Date: 2023-12-19 03:50 > >> > > To: gcc-patches > >> > > CC: rdapp.gcc; Li, Pan2; juzhe.zhong@rivai.ai > >> > > Subject: [PATCH] fold-const: Handle AND, IOR, XOR with stepped vec= tors [PR112971]. > >> > > Hi, > >> > > > >> > > found in PR112971, this patch adds folding support for bitwise ope= rations > >> > > of const duplicate zero vectors and stepped vectors. > >> > > On riscv we have the situation that a folding would perpetually co= ntinue > >> > > without simplifying because e.g. {0, 0, 0, ...} & {7, 6, 5, ...} w= ould > >> > > not fold to {0, 0, 0, ...}. > >> > > > >> > > Bootstrapped and regtested on x86 and aarch64, regtested on riscv. > >> > > > >> > > I won't be available to respond quickly until next year. Pan or J= uzhe, > >> > > as discussed, feel free to continue with possible revisions. > >> > > > >> > > Regards > >> > > Robin > >> > > > >> > > > >> > > gcc/ChangeLog: > >> > > > >> > > PR middle-end/112971 > >> > > > >> > > * fold-const.cc (const_binop): Handle > >> > > zerop@1 AND/IOR/XOR VECT_CST_STEPPED_P@2 > >> > > > >> > > gcc/testsuite/ChangeLog: > >> > > > >> > > * gcc.target/riscv/rvv/autovec/pr112971.c: New test. > >> > > --- > >> > > gcc/fold-const.cc | 14 +++++++++++++- > >> > > .../gcc.target/riscv/rvv/autovec/pr112971.c | 18 ++++++++++++++= ++++ > >> > > 2 files changed, 31 insertions(+), 1 deletion(-) > >> > > create mode 100644 gcc/testsuite/gcc.target/riscv/rvv/autovec/pr11= 2971.c > >> > > > >> > > diff --git a/gcc/fold-const.cc b/gcc/fold-const.cc > >> > > index f5d68ac323a..43ed097bf5c 100644 > >> > > --- a/gcc/fold-const.cc > >> > > +++ b/gcc/fold-const.cc > >> > > @@ -1653,8 +1653,20 @@ const_binop (enum tree_code code, tree arg1= , tree arg2) > >> > > { > >> > > tree type =3D TREE_TYPE (arg1); > >> > > bool step_ok_p; > >> > > + > >> > > + /* AND, IOR as well as XOR with a zerop can be handled dire= ctly. */ > >> > > if (VECTOR_CST_STEPPED_P (arg1) > >> > > - && VECTOR_CST_STEPPED_P (arg2)) > >> > > + && VECTOR_CST_DUPLICATE_P (arg2) > >> > > + && integer_zerop (VECTOR_CST_ELT (arg2, 0))) > >> > > + step_ok_p =3D code =3D=3D BIT_AND_EXPR || code =3D=3D BIT_IOR_EX= PR > >> > > + || code =3D=3D BIT_XOR_EXPR; > >> > > + else if (VECTOR_CST_STEPPED_P (arg2) > >> > > + && VECTOR_CST_DUPLICATE_P (arg1) > >> > > + && integer_zerop (VECTOR_CST_ELT (arg1, 0))) > >> > > + step_ok_p =3D code =3D=3D BIT_AND_EXPR || code =3D=3D BIT_IOR_EX= PR > >> > > + || code =3D=3D BIT_XOR_EXPR; > >> > > + else if (VECTOR_CST_STEPPED_P (arg1) > >> > > + && VECTOR_CST_STEPPED_P (arg2)) > >> > > /* We can operate directly on the encoding if: > >> > > a3 - a2 =3D=3D a2 - a1 && b3 - b2 =3D=3D b2 - b1 > >> > > diff --git a/gcc/testsuite/gcc.target/riscv/rvv/autovec/pr112971.c= b/gcc/testsuite/gcc.target/riscv/rvv/autovec/pr112971.c > >> > > new file mode 100644 > >> > > index 00000000000..816ebd3c493 > >> > > --- /dev/null > >> > > +++ b/gcc/testsuite/gcc.target/riscv/rvv/autovec/pr112971.c > >> > > @@ -0,0 +1,18 @@ > >> > > +/* { dg-do compile } */ > >> > > +/* { dg-options "-march=3Drv64gcv_zvl256b -mabi=3Dlp64d -O3 -fno-= vect-cost-model" } */ > >> > > + > >> > > +int a; > >> > > +short b[9]; > >> > > +char c, d; > >> > > +void e() { > >> > > + d =3D 0; > >> > > + for (;; d++) { > >> > > + if (b[d]) > >> > > + break; > >> > > + a =3D 8; > >> > > + for (; a >=3D 0; a--) { > >> > > + char *f =3D &c; > >> > > + *f &=3D d =3D=3D (a & d); > >> > > + } > >> > > + } > >> > > +} > >> > > > >> > > >> > > >> > >>