From mboxrd@z Thu Jan  1 00:00:00 1970
Return-Path: <philipp.tomsich@vrull.eu>
Received: from mail-lf1-x12a.google.com (mail-lf1-x12a.google.com [IPv6:2a00:1450:4864:20::12a])
	by sourceware.org (Postfix) with ESMTPS id B7CE43858D3C
	for <gcc-patches@gcc.gnu.org>; Mon, 14 Nov 2022 14:56:37 +0000 (GMT)
DMARC-Filter: OpenDMARC Filter v1.4.1 sourceware.org B7CE43858D3C
Authentication-Results: sourceware.org; dmarc=none (p=none dis=none) header.from=vrull.eu
Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=vrull.eu
Received: by mail-lf1-x12a.google.com with SMTP id j4so19757020lfk.0
        for <gcc-patches@gcc.gnu.org>; Mon, 14 Nov 2022 06:56:37 -0800 (PST)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=vrull.eu; s=google;
        h=content-transfer-encoding:cc:to:subject:message-id:date:from
         :in-reply-to:references:mime-version:from:to:cc:subject:date
         :message-id:reply-to;
        bh=mzq5XRdd/OzsUuJ40Vf/Ns4kKNPlJ3hxqFsvmS42qrQ=;
        b=TVaWzt1LE3pAr9APVLclGeI5wW0y8YnuOuYUgm+0pnClve34t2OzQlaIFlJ3E0jus2
         8UYPGkc0CSMC5lGPNDQBB6E5fPaGUtoUllrFVf4GZlolQ9n4zH4QuYts0fkNFU7S4RxO
         ZasaHdi9N5krMdetXJXrAahbPOv9prf5ZChWUj5YkZt6REBmqpNh2Clm6mEoNEpiLkpy
         ApkJw681aAn2nOkaaECtL0EEbMQCmgu0batTgliJsc5lUBFtpfAH9tW+O5OsavsboEOJ
         K9IuFQUHr4bDWEaTyoqyDqbwKUGBBqJ9ny/mvCzQIQs/KhtIsQgLRJUPqUHoFKb+AQF6
         is/g==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
        d=1e100.net; s=20210112;
        h=content-transfer-encoding:cc:to:subject:message-id:date:from
         :in-reply-to:references:mime-version:x-gm-message-state:from:to:cc
         :subject:date:message-id:reply-to;
        bh=mzq5XRdd/OzsUuJ40Vf/Ns4kKNPlJ3hxqFsvmS42qrQ=;
        b=hhD8o9qAisl1YGEZd8ltWJqkqGo0nOq9hnn2xljmjlBnA5rhoQK6FsGpsq32pi6CPy
         oNMMUCW18t28kr83/W8iTRIk9OZqiu8LjlKwya3VasojDiPmHHSrSYtw7ztvDRpQQsCt
         Gqr/RdA51l9gatBxIMR6Pc63CU6Pkq1ujAB63RKQNFRjs78UcgBcIeuY74FesTdhr19Y
         tir/1/sGKS3HXxtszsOtT28C7CbYOjMpCmY0/IzRU+/xDMavSLjniTQiv5+rQXnGsG8H
         euVnE4HtZt3imyXU+dLjBrAsL98uZfTxndCjGVxu8QNxLngg8Rtt0nubdfZo+VIpMStK
         irgQ==
X-Gm-Message-State: ANoB5plKx4OZ90TPlpAcc9p+X3M7XSRq4qDjyOGE/lkave9DQS3v1LZq
	PPmdYwJ5oGrjOjilJhdTuvzm/iV3GSNpvfsWzlls3NWwU+xfG5ul
X-Google-Smtp-Source: AA0mqf7IdQ5r9gD2oWh9SE880zK3IeXtOwuB7HjhY1MaU/CXSefTa1BB64DoiskawjXq9RcPGX3J220UiW+uWwBwFnE=
X-Received: by 2002:ac2:558c:0:b0:4a2:7692:3a0a with SMTP id
 v12-20020ac2558c000000b004a276923a0amr4449226lfg.71.1668437796243; Mon, 14
 Nov 2022 06:56:36 -0800 (PST)
MIME-Version: 1.0
References: <20221114135324.19352-1-philipp.tomsich@vrull.eu>
In-Reply-To: <20221114135324.19352-1-philipp.tomsich@vrull.eu>
From: Philipp Tomsich <philipp.tomsich@vrull.eu>
Date: Mon, 14 Nov 2022 15:56:25 +0100
Message-ID: <CAAeLtUCyj+oRbfufnGWPY+Xx1yCNQejGMB1Wo4myL4jqkL04Mw@mail.gmail.com>
Subject: Re: [PATCH v2] aarch64: Add support for Ampere-1A (-mcpu=ampere1a) CPU
To: Richard Sandiford <richard.sandiford@arm.com>
Cc: gcc-patches@gcc.gnu.org
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable
X-Spam-Status: No, score=-8.9 required=5.0 tests=BAYES_00,DKIM_SIGNED,DKIM_VALID,DKIM_VALID_AU,DKIM_VALID_EF,GIT_PATCH_0,JMQ_SPF_NEUTRAL,RCVD_IN_DNSWL_NONE,SPF_HELO_NONE,SPF_PASS,TXREP autolearn=ham autolearn_force=no version=3.4.6
X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on server2.sourceware.org
List-Id: <gcc-patches.gcc.gnu.org>

Richard,

is this OK for backport to GCC-12 and GCC-11?

Thanks,
Philipp.

On Mon, 14 Nov 2022 at 14:53, Philipp Tomsich <philipp.tomsich@vrull.eu> wr=
ote:
>
> This patch adds support for Ampere-1A CPU:
>  - recognize the name of the core and provide detection for -mcpu=3Dnativ=
e,
>  - updated extra_costs,
>  - adds a new fusion pair for (A+B+1 and A-B-1).
>
> Ampere-1A and Ampere-1 have more timing difference than the extra
> costs indicate, but these don't propagate through to the headline
> items in our extra costs (e.g. the change in latency for scalar sqrt
> doesn't have a corresponding table entry).
>
> gcc/ChangeLog:
>
>         * config/aarch64/aarch64-cores.def (AARCH64_CORE): Add ampere1a.
>         * config/aarch64/aarch64-cost-tables.h: Add ampere1a_extra_costs.
>         * config/aarch64/aarch64-fusion-pairs.def (AARCH64_FUSION_PAIR):
>         Define a new fusion pair for A+B+1/A-B-1 (i.e., add/subtract two
>         registers and then +1/-1).
>         * config/aarch64/aarch64-tune.md: Regenerate.
>         * config/aarch64/aarch64.cc (aarch_macro_fusion_pair_p): Implemen=
t
>         idiom-matcher for the new fusion pair.
>         * doc/invoke.texi: Add ampere1a.
>
> Signed-off-by: Philipp Tomsich <philipp.tomsich@vrull.eu>
> ---
>
> Changes in v2:
> - break line in fusion matcher to stay below 80 characters
> - rename fusion pair addsub_2reg_const1
> - document 'ampere1a' in invoke.texi
>
>  gcc/config/aarch64/aarch64-cores.def        |   1 +
>  gcc/config/aarch64/aarch64-cost-tables.h    | 107 ++++++++++++++++++++
>  gcc/config/aarch64/aarch64-fusion-pairs.def |   1 +
>  gcc/config/aarch64/aarch64-tune.md          |   2 +-
>  gcc/config/aarch64/aarch64.cc               |  64 ++++++++++++
>  gcc/doc/invoke.texi                         |   2 +-
>  6 files changed, 175 insertions(+), 2 deletions(-)
>
> diff --git a/gcc/config/aarch64/aarch64-cores.def b/gcc/config/aarch64/aa=
rch64-cores.def
> index d2671778928..aead587cec1 100644
> --- a/gcc/config/aarch64/aarch64-cores.def
> +++ b/gcc/config/aarch64/aarch64-cores.def
> @@ -70,6 +70,7 @@ AARCH64_CORE("thunderxt83",   thunderxt83,   thunderx, =
 V8A,  (CRC, CRYPTO), thu
>
>  /* Ampere Computing ('\xC0') cores. */
>  AARCH64_CORE("ampere1", ampere1, cortexa57, V8_6A, (F16, RNG, AES, SHA3)=
, ampere1, 0xC0, 0xac3, -1)
> +AARCH64_CORE("ampere1a", ampere1a, cortexa57, V8_6A, (F16, RNG, AES, SHA=
3, MEMTAG), ampere1a, 0xC0, 0xac4, -1)
>  /* Do not swap around "emag" and "xgene1",
>     this order is required to handle variant correctly. */
>  AARCH64_CORE("emag",        emag,      xgene1,    V8A,  (CRC, CRYPTO), e=
mag, 0x50, 0x000, 3)
> diff --git a/gcc/config/aarch64/aarch64-cost-tables.h b/gcc/config/aarch6=
4/aarch64-cost-tables.h
> index 760d7b30368..48522606fbe 100644
> --- a/gcc/config/aarch64/aarch64-cost-tables.h
> +++ b/gcc/config/aarch64/aarch64-cost-tables.h
> @@ -775,4 +775,111 @@ const struct cpu_cost_table ampere1_extra_costs =3D
>    }
>  };
>
> +const struct cpu_cost_table ampere1a_extra_costs =3D
> +{
> +  /* ALU */
> +  {
> +    0,                 /* arith.  */
> +    0,                 /* logical.  */
> +    0,                 /* shift.  */
> +    COSTS_N_INSNS (1), /* shift_reg.  */
> +    0,                 /* arith_shift.  */
> +    COSTS_N_INSNS (1), /* arith_shift_reg.  */
> +    0,                 /* log_shift.  */
> +    COSTS_N_INSNS (1), /* log_shift_reg.  */
> +    0,                 /* extend.  */
> +    COSTS_N_INSNS (1), /* extend_arith.  */
> +    0,                 /* bfi.  */
> +    0,                 /* bfx.  */
> +    0,                 /* clz.  */
> +    0,                 /* rev.  */
> +    0,                 /* non_exec.  */
> +    true               /* non_exec_costs_exec.  */
> +  },
> +  {
> +    /* MULT SImode */
> +    {
> +      COSTS_N_INSNS (3),       /* simple.  */
> +      COSTS_N_INSNS (3),       /* flag_setting.  */
> +      COSTS_N_INSNS (3),       /* extend.  */
> +      COSTS_N_INSNS (4),       /* add.  */
> +      COSTS_N_INSNS (4),       /* extend_add.  */
> +      COSTS_N_INSNS (19)       /* idiv.  */
> +    },
> +    /* MULT DImode */
> +    {
> +      COSTS_N_INSNS (3),       /* simple.  */
> +      0,                       /* flag_setting (N/A).  */
> +      COSTS_N_INSNS (3),       /* extend.  */
> +      COSTS_N_INSNS (4),       /* add.  */
> +      COSTS_N_INSNS (4),       /* extend_add.  */
> +      COSTS_N_INSNS (35)       /* idiv.  */
> +    }
> +  },
> +  /* LD/ST */
> +  {
> +    COSTS_N_INSNS (4),         /* load.  */
> +    COSTS_N_INSNS (4),         /* load_sign_extend.  */
> +    0,                         /* ldrd (n/a).  */
> +    0,                         /* ldm_1st.  */
> +    0,                         /* ldm_regs_per_insn_1st.  */
> +    0,                         /* ldm_regs_per_insn_subsequent.  */
> +    COSTS_N_INSNS (5),         /* loadf.  */
> +    COSTS_N_INSNS (5),         /* loadd.  */
> +    COSTS_N_INSNS (5),         /* load_unaligned.  */
> +    0,                         /* store.  */
> +    0,                         /* strd.  */
> +    0,                         /* stm_1st.  */
> +    0,                         /* stm_regs_per_insn_1st.  */
> +    0,                         /* stm_regs_per_insn_subsequent.  */
> +    COSTS_N_INSNS (2),         /* storef.  */
> +    COSTS_N_INSNS (2),         /* stored.  */
> +    COSTS_N_INSNS (2),         /* store_unaligned.  */
> +    COSTS_N_INSNS (3),         /* loadv.  */
> +    COSTS_N_INSNS (3)          /* storev.  */
> +  },
> +  {
> +    /* FP SFmode */
> +    {
> +      COSTS_N_INSNS (25),      /* div.  */
> +      COSTS_N_INSNS (4),       /* mult.  */
> +      COSTS_N_INSNS (4),       /* mult_addsub.  */
> +      COSTS_N_INSNS (4),       /* fma.  */
> +      COSTS_N_INSNS (4),       /* addsub.  */
> +      COSTS_N_INSNS (2),       /* fpconst.  */
> +      COSTS_N_INSNS (4),       /* neg.  */
> +      COSTS_N_INSNS (4),       /* compare.  */
> +      COSTS_N_INSNS (4),       /* widen.  */
> +      COSTS_N_INSNS (4),       /* narrow.  */
> +      COSTS_N_INSNS (4),       /* toint.  */
> +      COSTS_N_INSNS (4),       /* fromint.  */
> +      COSTS_N_INSNS (4)        /* roundint.  */
> +    },
> +    /* FP DFmode */
> +    {
> +      COSTS_N_INSNS (34),      /* div.  */
> +      COSTS_N_INSNS (5),       /* mult.  */
> +      COSTS_N_INSNS (5),       /* mult_addsub.  */
> +      COSTS_N_INSNS (5),       /* fma.  */
> +      COSTS_N_INSNS (5),       /* addsub.  */
> +      COSTS_N_INSNS (2),       /* fpconst.  */
> +      COSTS_N_INSNS (5),       /* neg.  */
> +      COSTS_N_INSNS (5),       /* compare.  */
> +      COSTS_N_INSNS (5),       /* widen.  */
> +      COSTS_N_INSNS (5),       /* narrow.  */
> +      COSTS_N_INSNS (6),       /* toint.  */
> +      COSTS_N_INSNS (6),       /* fromint.  */
> +      COSTS_N_INSNS (5)        /* roundint.  */
> +    }
> +  },
> +  /* Vector */
> +  {
> +    COSTS_N_INSNS (3),  /* alu.  */
> +    COSTS_N_INSNS (3),  /* mult.  */
> +    COSTS_N_INSNS (2),  /* movi.  */
> +    COSTS_N_INSNS (2),  /* dup.  */
> +    COSTS_N_INSNS (2)   /* extract.  */
> +  }
> +};
> +
>  #endif
> diff --git a/gcc/config/aarch64/aarch64-fusion-pairs.def b/gcc/config/aar=
ch64/aarch64-fusion-pairs.def
> index c064fb9b85d..d91f8a2babd 100644
> --- a/gcc/config/aarch64/aarch64-fusion-pairs.def
> +++ b/gcc/config/aarch64/aarch64-fusion-pairs.def
> @@ -36,5 +36,6 @@ AARCH64_FUSION_PAIR ("cmp+branch", CMP_BRANCH)
>  AARCH64_FUSION_PAIR ("aes+aesmc", AES_AESMC)
>  AARCH64_FUSION_PAIR ("alu+branch", ALU_BRANCH)
>  AARCH64_FUSION_PAIR ("alu+cbz", ALU_CBZ)
> +AARCH64_FUSION_PAIR ("addsub_2reg_const1", ADDSUB_2REG_CONST1)
>
>  #undef AARCH64_FUSION_PAIR
> diff --git a/gcc/config/aarch64/aarch64-tune.md b/gcc/config/aarch64/aarc=
h64-tune.md
> index 22ec1be5a4c..b7d6fc8cc88 100644
> --- a/gcc/config/aarch64/aarch64-tune.md
> +++ b/gcc/config/aarch64/aarch64-tune.md
> @@ -1,5 +1,5 @@
>  ;; -*- buffer-read-only: t -*-
>  ;; Generated automatically by gentune.sh from aarch64-cores.def
>  (define_attr "tune"
> -       "cortexa34,cortexa35,cortexa53,cortexa57,cortexa72,cortexa73,thun=
derx,thunderxt88p1,thunderxt88,octeontx,octeontxt81,octeontxt83,thunderxt81=
,thunderxt83,ampere1,emag,xgene1,falkor,qdf24xx,exynosm1,phecda,thunderx2t9=
9p1,vulcan,thunderx2t99,cortexa55,cortexa75,cortexa76,cortexa76ae,cortexa77=
,cortexa78,cortexa78ae,cortexa78c,cortexa65,cortexa65ae,cortexx1,cortexx1c,=
ares,neoversen1,neoversee1,octeontx2,octeontx2t98,octeontx2t96,octeontx2t93=
,octeontx2f95,octeontx2f95n,octeontx2f95mm,a64fx,tsv110,thunderx3t110,zeus,=
neoversev1,neoverse512tvb,saphira,cortexa57cortexa53,cortexa72cortexa53,cor=
texa73cortexa35,cortexa73cortexa53,cortexa75cortexa55,cortexa76cortexa55,co=
rtexr82,cortexa510,cortexa710,cortexa715,cortexx2,neoversen2,demeter,neover=
sev2"
> +       "cortexa34,cortexa35,cortexa53,cortexa57,cortexa72,cortexa73,thun=
derx,thunderxt88p1,thunderxt88,octeontx,octeontxt81,octeontxt83,thunderxt81=
,thunderxt83,ampere1,ampere1a,emag,xgene1,falkor,qdf24xx,exynosm1,phecda,th=
underx2t99p1,vulcan,thunderx2t99,cortexa55,cortexa75,cortexa76,cortexa76ae,=
cortexa77,cortexa78,cortexa78ae,cortexa78c,cortexa65,cortexa65ae,cortexx1,c=
ortexx1c,ares,neoversen1,neoversee1,octeontx2,octeontx2t98,octeontx2t96,oct=
eontx2t93,octeontx2f95,octeontx2f95n,octeontx2f95mm,a64fx,tsv110,thunderx3t=
110,zeus,neoversev1,neoverse512tvb,saphira,cortexa57cortexa53,cortexa72cort=
exa53,cortexa73cortexa35,cortexa73cortexa53,cortexa75cortexa55,cortexa76cor=
texa55,cortexr82,cortexa510,cortexa710,cortexa715,cortexx2,neoversen2,demet=
er,neoversev2"
>         (const (symbol_ref "((enum attr_tune) aarch64_tune)")))
> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.c=
c
> index d1f979ebcf8..a7f7c3c0121 100644
> --- a/gcc/config/aarch64/aarch64.cc
> +++ b/gcc/config/aarch64/aarch64.cc
> @@ -1921,6 +1921,43 @@ static const struct tune_params ampere1_tunings =
=3D
>    &ampere1_prefetch_tune
>  };
>
> +static const struct tune_params ampere1a_tunings =3D
> +{
> +  &ampere1a_extra_costs,
> +  &generic_addrcost_table,
> +  &generic_regmove_cost,
> +  &ampere1_vector_cost,
> +  &generic_branch_cost,
> +  &generic_approx_modes,
> +  SVE_NOT_IMPLEMENTED, /* sve_width  */
> +  { 4, /* load_int.  */
> +    4, /* store_int.  */
> +    4, /* load_fp.  */
> +    4, /* store_fp.  */
> +    4, /* load_pred.  */
> +    4 /* store_pred.  */
> +  }, /* memmov_cost.  */
> +  4, /* issue_rate  */
> +  (AARCH64_FUSE_ADRP_ADD | AARCH64_FUSE_AES_AESMC |
> +   AARCH64_FUSE_MOV_MOVK | AARCH64_FUSE_MOVK_MOVK |
> +   AARCH64_FUSE_ALU_BRANCH /* adds, ands, bics, ccmp, ccmn */ |
> +   AARCH64_FUSE_CMP_BRANCH | AARCH64_FUSE_ALU_CBZ |
> +   AARCH64_FUSE_ADDSUB_2REG_CONST1),
> +  /* fusible_ops  */
> +  "32",                /* function_align.  */
> +  "4",         /* jump_align.  */
> +  "32:16",     /* loop_align.  */
> +  2,   /* int_reassoc_width.  */
> +  4,   /* fp_reassoc_width.  */
> +  2,   /* vec_reassoc_width.  */
> +  2,   /* min_div_recip_mul_sf.  */
> +  2,   /* min_div_recip_mul_df.  */
> +  0,   /* max_case_values.  */
> +  tune_params::AUTOPREFETCHER_WEAK,    /* autoprefetcher_model.  */
> +  (AARCH64_EXTRA_TUNE_NONE),           /* tune_flags.  */
> +  &ampere1_prefetch_tune
> +};
> +
>  static const advsimd_vec_cost neoversev1_advsimd_vector_cost =3D
>  {
>    2, /* int_stmt_cost  */
> @@ -25539,6 +25576,33 @@ aarch_macro_fusion_pair_p (rtx_insn *prev, rtx_i=
nsn *curr)
>         }
>      }
>
> +  /* Fuse A+B+1 and A-B-1 */
> +  if (simple_sets_p
> +      && aarch64_fusion_enabled_p (AARCH64_FUSE_ADDSUB_2REG_CONST1))
> +    {
> +      /* We're trying to match:
> +         prev =3D=3D (set (r0) (plus (r0) (r1)))
> +         curr =3D=3D (set (r0) (plus (r0) (const_int 1)))
> +       or:
> +         prev =3D=3D (set (r0) (minus (r0) (r1)))
> +         curr =3D=3D (set (r0) (plus (r0) (const_int -1))) */
> +
> +      rtx prev_src =3D SET_SRC (prev_set);
> +      rtx curr_src =3D SET_SRC (curr_set);
> +
> +      int polarity =3D 1;
> +      if (GET_CODE (prev_src) =3D=3D MINUS)
> +       polarity =3D -1;
> +
> +      if (GET_CODE (curr_src) =3D=3D PLUS
> +         && (GET_CODE (prev_src) =3D=3D PLUS || GET_CODE (prev_src) =3D=
=3D MINUS)
> +         && CONST_INT_P (XEXP (curr_src, 1))
> +         && INTVAL (XEXP (curr_src, 1)) =3D=3D polarity
> +         && REG_P (XEXP (curr_src, 0))
> +         && REGNO (SET_DEST (prev_set)) =3D=3D REGNO (XEXP (curr_src, 0)=
))
> +       return true;
> +    }
> +
>    return false;
>  }
>
> diff --git a/gcc/doc/invoke.texi b/gcc/doc/invoke.texi
> index 60e65f4eaa5..09c8b312ae7 100644
> --- a/gcc/doc/invoke.texi
> +++ b/gcc/doc/invoke.texi
> @@ -19995,7 +19995,7 @@ performance of the code.  Permissible values for =
this option are:
>  @samp{cortex-a75.cortex-a55}, @samp{cortex-a76.cortex-a55},
>  @samp{cortex-r82}, @samp{cortex-x1}, @samp{cortex-x1c}, @samp{cortex-x2}=
,
>  @samp{cortex-a510}, @samp{cortex-a710}, @samp{cortex-a715}, @samp{ampere=
1},
> -@samp{native}.
> +@samp{ampere1a}, and @samp{native}.
>
>  The values @samp{cortex-a57.cortex-a53}, @samp{cortex-a72.cortex-a53},
>  @samp{cortex-a73.cortex-a35}, @samp{cortex-a73.cortex-a53},
> --
> 2.34.1
>